<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DevOps Daily</title>
    <description>The latest articles on DEV Community by DevOps Daily (@devopsdaily).</description>
    <link>https://dev.to/devopsdaily</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F382434%2F66e04ef9-7e4f-491c-997a-30f4a999d40c.jpg</url>
      <title>DEV Community: DevOps Daily</title>
      <link>https://dev.to/devopsdaily</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devopsdaily"/>
    <language>en</language>
    <item>
      <title>How to resolve merge conflicts in Git without guessing: markers, --ours/--theirs and the rebase flip</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Thu, 08 Oct 2026 10:15:49 +0000</pubDate>
      <link>https://dev.to/devopsdaily/how-to-resolve-merge-conflicts-in-git-without-guessing-markers-ours-theirs-and-the-rebase-flip-b4g</link>
      <guid>https://dev.to/devopsdaily/how-to-resolve-merge-conflicts-in-git-without-guessing-markers-ours-theirs-and-the-rebase-flip-b4g</guid>
      <description>&lt;p&gt;You are rebasing your branch onto main. Git stops on a conflict in &lt;code&gt;limits.py&lt;/code&gt;. You want to keep your version, so you type &lt;code&gt;git checkout --ours limits.py&lt;/code&gt;, stage it, continue and push. Your change is gone. During a rebase, "ours" is main, and "theirs" is the commit you wrote.&lt;/p&gt;

&lt;p&gt;That one swap is why so many people learn how to resolve merge conflicts in Git by trial and error: keep the top half, run the tests, try the bottom half. Most tutorials stop at "delete the markers and commit". This one covers the parts that bite later, including the rebase flip and the clean merge that still breaks the build.&lt;/p&gt;

&lt;p&gt;The hands-on part is the &lt;a href="https://devops-daily.com/games/git-merge-conflict-simulator" rel="noopener noreferrer"&gt;Git Merge Conflict Simulator&lt;/a&gt;: six lessons and 27 steps. Five open on a repository that is already mid-conflict; the sixth starts just before a clean merge that breaks the code. You type commands into a terminal, edit files in an editor, and watch the "Index stages" and "Ours and theirs" panels change. It does not run real Git: it models the commands each lesson needs, with output based on real Git's. Disclosure: I help build it; it is free, runs in the browser, no signup.&lt;/p&gt;

&lt;h2&gt;
  
  
  A conflict is three versions, not two
&lt;/h2&gt;

&lt;p&gt;Lesson 1, &lt;strong&gt;Read the conflict&lt;/strong&gt;, opens after &lt;code&gt;git merge feature/retry&lt;/code&gt; has stopped. &lt;code&gt;config.yaml&lt;/code&gt; started with &lt;code&gt;timeout: 30&lt;/code&gt;. Main raised it to 45, the feature branch to 60. The markers show two of the three versions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;service: api
&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt; HEAD
timeout: 45
=======
timeout: 60
&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt; feature/retry
retries: 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The third version, the common ancestor, is the most useful one and the one the markers hide. While a path is unmerged, Git keeps up to three copies of it in the index: stage 1 is the base, stage 2 is ours, stage 3 is theirs. The lesson has you run &lt;code&gt;git show :1:config.yaml&lt;/code&gt; before you touch anything. Once you know the base was 30, you know both sides raised the timeout, so the question is "which increase wins", not "who broke it".&lt;/p&gt;

&lt;p&gt;If you want the base inline every time, run &lt;code&gt;git config --global merge.conflictStyle zdiff3&lt;/code&gt; (added in Git 2.35). Conflicts then show a &lt;code&gt;|||||||&lt;/code&gt; section with the base between the two sides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resolving means writing the file you want
&lt;/h2&gt;

&lt;p&gt;Lesson 2, &lt;strong&gt;Resolve by editing&lt;/strong&gt;: main added &lt;code&gt;redis==5.0.8&lt;/code&gt; to &lt;code&gt;requirements.txt&lt;/code&gt;, and a metrics branch added &lt;code&gt;prometheus-client==0.20.0&lt;/code&gt; at the same spot. The project needs both. Neither &lt;code&gt;--ours&lt;/code&gt; nor &lt;code&gt;--theirs&lt;/code&gt; gives you that, so you open the editor, delete the three marker lines, keep both packages and save. A save that still has markers or drops a package does not complete the step.&lt;/p&gt;

&lt;p&gt;Then &lt;code&gt;git add requirements.txt&lt;/code&gt;. That is how you tell Git a conflict is resolved: it replaces the three index stages with the one version you staged. Git does not read the content, so it will stage a file full of markers without complaint. In real Git, &lt;code&gt;git diff --cached --check&lt;/code&gt; reports "leftover conflict marker" lines before you commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Take a whole side when nobody can merge it by hand
&lt;/h2&gt;

&lt;p&gt;Lesson 3, &lt;strong&gt;Take a whole side&lt;/strong&gt;: &lt;code&gt;assets/logo.png&lt;/code&gt; is binary, so Git writes no markers at all. The working tree keeps main's logo and the index holds all three versions. The merge is a rebrand, so you want the incoming file: &lt;code&gt;git checkout --theirs assets/logo.png&lt;/code&gt;, then &lt;code&gt;git add&lt;/code&gt; and &lt;code&gt;git commit --no-edit&lt;/code&gt;. During a merge, &lt;code&gt;--ours&lt;/code&gt; is the branch you are on and &lt;code&gt;--theirs&lt;/code&gt; is the branch you are merging in. Taking a side copies the whole file, so main's sharper outline is lost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rebase flip
&lt;/h2&gt;

&lt;p&gt;Lesson 4, &lt;strong&gt;Rebase flips ours and theirs&lt;/strong&gt;, is the opening story. A rebase checks out main's tip (HEAD is detached there) and replays your commits on top of it, one at a time. So &lt;code&gt;HEAD&lt;/code&gt; and &lt;code&gt;--ours&lt;/code&gt; mean main plus whatever of yours has been replayed so far, and &lt;code&gt;--theirs&lt;/code&gt; is your commit being replayed. The simulator's side panel shows both rules:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;During a&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;--ours&lt;/code&gt; / &lt;code&gt;HEAD&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;&lt;code&gt;--theirs&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;merge&lt;/td&gt;
&lt;td&gt;the branch you are on&lt;/td&gt;
&lt;td&gt;the branch you are merging in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rebase&lt;/td&gt;
&lt;td&gt;the rebased result so far&lt;/td&gt;
&lt;td&gt;your commit being replayed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;git checkout&lt;/code&gt; man page says the same: during &lt;code&gt;git rebase&lt;/code&gt; and &lt;code&gt;git pull --rebase&lt;/code&gt;, ours and theirs "may appear swapped". In the lesson the team agreed your stricter limit of 30 requests per minute beats main's 200, so the right command is &lt;code&gt;git checkout --theirs limits.py&lt;/code&gt;, then &lt;code&gt;git add&lt;/code&gt; and &lt;code&gt;git rebase --continue&lt;/code&gt;. If you already typed &lt;code&gt;--ours&lt;/code&gt; and have not staged it, &lt;code&gt;git checkout -m limits.py&lt;/code&gt; recreates the conflict.&lt;/p&gt;

&lt;h2&gt;
  
  
  The escape hatch
&lt;/h2&gt;

&lt;p&gt;Lesson 5, &lt;strong&gt;The escape hatch&lt;/strong&gt;: you merged &lt;code&gt;release/2025.12&lt;/code&gt; when you meant &lt;code&gt;release/2026.09&lt;/code&gt;, and six files are conflicted. None of them is worth resolving. &lt;code&gt;git merge --abort&lt;/code&gt; tries to put the index and working tree back and removes &lt;code&gt;MERGE_HEAD&lt;/code&gt;; &lt;code&gt;git rebase --abort&lt;/code&gt; does the same for a rebase. The caveat the lesson spells out: if you had uncommitted changes when you started the merge, &lt;code&gt;--abort&lt;/code&gt; may not be able to restore them. Commit or stash first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Clean merge, broken code
&lt;/h2&gt;

&lt;p&gt;Lesson 6, &lt;strong&gt;Clean merge, broken code&lt;/strong&gt;, has no markers at all. Main renamed &lt;code&gt;settings.get_timeout()&lt;/code&gt; to &lt;code&gt;get_request_timeout()&lt;/code&gt; and fixed every caller it had. A week-old branch adds a new caller of the old name in a different file. &lt;code&gt;git merge feature/worker&lt;/code&gt; succeeds without asking anything, &lt;code&gt;make test&lt;/code&gt; fails, and you fix one line in &lt;code&gt;worker.py&lt;/code&gt;. Each branch passed its own tests; only the combination is broken. Git merges text, not behaviour, which is why CI should run on the merge result, or a merge queue should test the combined code before it lands.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the flip in a real repo
&lt;/h2&gt;

&lt;p&gt;To see lesson 4 on your own machine, build the same situation in a scratch directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;conflict-lab &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;conflict-lab
git init &lt;span class="nt"&gt;-b&lt;/span&gt; main
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"limit = 100"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; limits.py
git add limits.py &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"Add rate limiting"&lt;/span&gt;

git switch &lt;span class="nt"&gt;-c&lt;/span&gt; feature/rate-limit
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"limit = 30"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; limits.py
git commit &lt;span class="nt"&gt;-am&lt;/span&gt; &lt;span class="s2"&gt;"Tighten the default rate limit"&lt;/span&gt;

git switch main
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"limit = 200"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; limits.py
git commit &lt;span class="nt"&gt;-am&lt;/span&gt; &lt;span class="s2"&gt;"Allow 200 requests per minute"&lt;/span&gt;

git switch feature/rate-limit
git rebase main
&lt;span class="nb"&gt;cat &lt;/span&gt;limits.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With the default conflict style on Git 2.39, the rebase stops and &lt;code&gt;cat&lt;/code&gt; prints this (your hash will differ, and with &lt;code&gt;zdiff3&lt;/code&gt; set you also get a &lt;code&gt;|||||||&lt;/code&gt; base section):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt; HEAD
limit = 200
=======
limit = 30
&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt; f9ce8c6 (Tighten the default rate limit)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The half labelled &lt;code&gt;HEAD&lt;/code&gt; is main's line, even though you started on your feature branch. &lt;code&gt;git show :2:limits.py&lt;/code&gt; prints &lt;code&gt;limit = 200&lt;/code&gt; and &lt;code&gt;git show :3:limits.py&lt;/code&gt; prints &lt;code&gt;limit = 30&lt;/code&gt;. Run &lt;code&gt;git checkout --theirs limits.py&lt;/code&gt;, &lt;code&gt;git add limits.py&lt;/code&gt; and &lt;code&gt;git rebase --continue&lt;/code&gt; (save the commit message when the editor opens), and your 30 survives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to go next
&lt;/h2&gt;

&lt;p&gt;Do the six lessons in order: the first three build the model that makes the fourth obvious. The simulator is one of &lt;a href="https://devops-daily.com/games" rel="noopener noreferrer"&gt;50+ free DevOps games and simulators&lt;/a&gt;. Official references: the "How conflicts are presented" section of &lt;a href="https://git-scm.com/docs/git-merge#_how_conflicts_are_presented" rel="noopener noreferrer"&gt;git-merge&lt;/a&gt; and the &lt;code&gt;--ours&lt;/code&gt;/&lt;code&gt;--theirs&lt;/code&gt; note in &lt;a href="https://git-scm.com/docs/git-checkout" rel="noopener noreferrer"&gt;git-checkout&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>git</category>
      <category>beginners</category>
      <category>tutorial</category>
      <category>productivity</category>
    </item>
    <item>
      <title>OpenTofu 1.13 Re-Encodes base64gzip and Drops WinRM. We Upgraded the Same State to See What Breaks</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Wed, 07 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/opentofu-113-re-encodes-base64gzip-and-drops-winrm-we-upgraded-the-same-state-to-see-what-breaks-12bh</link>
      <guid>https://dev.to/devopsdaily/opentofu-113-re-encodes-base64gzip-and-drops-winrm-we-upgraded-the-same-state-to-see-what-breaks-12bh</guid>
      <description>&lt;p&gt;OpenTofu 1.13.0 was published on September 30 (the &lt;a href="https://opentofu.org/blog/opentofu-1-13-0/" rel="noopener noreferrer"&gt;release announcement&lt;/a&gt; is dated September 29), and &lt;a href="https://github.com/opentofu/opentofu/releases/tag/v1.13.1" rel="noopener noreferrer"&gt;1.13.1&lt;/a&gt; followed on October 1 with two fixes for ephemeral values. The headline features are new functions that let module authors describe values OpenTofu cannot know until apply, plus two experiments: built-in linting and symbol libraries. The upgrade notes are what you will notice first. &lt;code&gt;base64gzip&lt;/code&gt; now returns different bytes for the same input, WinRM provisioners are gone, and 1.13 is the last series with official 32-bit builds.&lt;/p&gt;

&lt;p&gt;Upgrade notes tell you what changed. They do not show you your next plan, or the step at which a removed feature fails. So we tested it. We applied configurations with OpenTofu 1.12.7, ran 1.13.1 against the same state, and recorded what each command printed. This post goes through the results and ends with a checklist you can run before you change the version in CI.&lt;/p&gt;

&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;base64gzip&lt;/code&gt; in 1.13.1 returns a different string for the same input. OpenTofu moved to Go 1.27, and Go changed its DEFLATE encoder. Both outputs decompress to identical bytes, but providers compare the string. With no config change, our plan was &lt;code&gt;1 to add, 1 to change, 1 to destroy&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;On &lt;code&gt;aws_instance&lt;/code&gt;, a changed &lt;code&gt;user_data_base64&lt;/code&gt; means a stop/start, or a replacement if you set &lt;code&gt;user_data_replace_on_change&lt;/code&gt;. On &lt;code&gt;azurerm_linux_virtual_machine&lt;/code&gt;, a changed &lt;code&gt;custom_data&lt;/code&gt; forces a new VM.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ignore_changes&lt;/code&gt; hides the diff. In our test it also hid a real edit to the cloud-init file. Gzip done inside the &lt;code&gt;cloudinit&lt;/code&gt; provider did not change at all.&lt;/li&gt;
&lt;li&gt;A &lt;code&gt;winrm&lt;/code&gt; provisioner passes &lt;code&gt;tofu validate&lt;/code&gt; and &lt;code&gt;tofu plan&lt;/code&gt; on 1.13.1, with the old "will be removed in a future version" warning. It fails at apply, after the resource exists, and leaves the resource tainted.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;assumenotnull&lt;/code&gt; fixed a classic &lt;code&gt;Invalid count argument&lt;/code&gt; error. &lt;code&gt;assumestringprefix&lt;/code&gt; moved a mis-wired module input from a mid-apply failure to a plan error. Older versions reject these functions, and since 1.12 a &lt;code&gt;required_version&lt;/code&gt; in a &lt;code&gt;.tf&lt;/code&gt; file does not stop them.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;-lint=all&lt;/code&gt; is a useful experiment, but it only adds warnings. The exit code stays 0.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;An OpenTofu 1.12.x codebase, or a Terraform codebase that you plan to move to OpenTofu&lt;/li&gt;
&lt;li&gt;A CI job or shell that can run &lt;code&gt;tofu plan&lt;/code&gt; with read access to your real state&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;jq&lt;/code&gt;, for the plan JSON recipe&lt;/li&gt;
&lt;li&gt;The 1.13.1 zip for your platform from the &lt;a href="https://github.com/opentofu/opentofu/releases/tag/v1.13.1" rel="noopener noreferrer"&gt;GitHub release page&lt;/a&gt;, checked against its &lt;code&gt;SHA256SUMS&lt;/code&gt; file&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How we tested
&lt;/h2&gt;

&lt;p&gt;We ran everything on a Raspberry Pi with a 64-bit (arm64) OS: OpenTofu 1.12.7 and 1.13.1 for &lt;code&gt;linux_arm64&lt;/code&gt;, the 1.13.1 &lt;code&gt;linux_arm&lt;/code&gt; (32-bit) build, and Terraform 1.16.5 for comparison. Every zip matched its published SHA256 sum. In the outputs below, &lt;code&gt;tofu-1.12.7&lt;/code&gt; and &lt;code&gt;tofu-1.13.1&lt;/code&gt; are the two release binaries side by side.&lt;/p&gt;

&lt;p&gt;We did not use a cloud account. The built-in &lt;code&gt;terraform_data&lt;/code&gt; resource stands in for provider attributes. Its &lt;code&gt;input&lt;/code&gt; argument updates in place when the value changes, and its &lt;code&gt;triggers_replace&lt;/code&gt; argument forces a replacement. Those are the two ways real providers treat user data, which we checked in the provider docs. Every output in this post comes from these runs. Where we cut lines from an output, the block shows &lt;code&gt;...&lt;/code&gt; or the &lt;code&gt;grep&lt;/code&gt;/&lt;code&gt;tail&lt;/code&gt; we used, and long base64 strings are shortened with &lt;code&gt;...&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  base64gzip: same input, different string
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;base64gzip&lt;/code&gt; compresses a string with gzip and base64-encodes the result. Its most common job is cloud-init user data, because EC2 limits user data to 16 KB and compression gives you more room. Here is the same call on both versions, plus Terraform 1.16.5, and then a check of what a real payload (a 1,208-byte cloud-init file that installs nginx and writes a systemd unit) decompresses to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;base64gzip across versions&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# same input, three binaries&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'base64gzip("hello")'&lt;/span&gt; | ./tofu-1.12.7 console
&lt;span class="s2"&gt;"H4sIAAAAAAAA/8pIzcnJBwAAAP//AQAA//+GphA2BQAAAA=="&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'base64gzip("hello")'&lt;/span&gt; | ./tofu-1.13.1 console
&lt;span class="s2"&gt;"H4sIAAAAAAAA/wAFAPr/aGVsbG8AAAD//wMAhqYQNgUAAAA="&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'base64gzip("hello")'&lt;/span&gt; | ./terraform-1.16.5 console
&lt;span class="s2"&gt;"H4sIAAAAAAAA/8pIzcnJBwAAAP//AQAA//+GphA2BQAAAA=="&lt;/span&gt;
&lt;span class="c"&gt;# decompress the cloud-init payload from each version&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'base64gzip(file("cloud-init.yaml"))'&lt;/span&gt; | ./tofu-1.12.7 console | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'"'&lt;/span&gt; | &lt;span class="nb"&gt;base64&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; | &lt;span class="nb"&gt;gunzip&lt;/span&gt; | &lt;span class="nb"&gt;sha256sum
&lt;/span&gt;3a62ee85006428492dc1aeccc789867735496026ada3bfaef3d0cf00dcc1bcd1 -
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'base64gzip(file("cloud-init.yaml"))'&lt;/span&gt; | ./tofu-1.13.1 console | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'"'&lt;/span&gt; | &lt;span class="nb"&gt;base64&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; | &lt;span class="nb"&gt;gunzip&lt;/span&gt; | &lt;span class="nb"&gt;sha256sum
&lt;/span&gt;3a62ee85006428492dc1aeccc789867735496026ada3bfaef3d0cf00dcc1bcd1 -
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sha256sum&lt;/span&gt; &amp;lt; cloud-init.yaml
3a62ee85006428492dc1aeccc789867735496026ada3bfaef3d0cf00dcc1bcd1 -

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The encoded strings differ. The content does not: both outputs decompress to the exact bytes of the source file.&lt;/p&gt;

&lt;p&gt;The cause is upstream. OpenTofu 1.12.7 is built with Go 1.26.6 and 1.13.1 with Go 1.27.1 (each binary records its Go version). The &lt;a href="https://go.dev/doc/go1.27" rel="noopener noreferrer"&gt;Go 1.27 release notes&lt;/a&gt; say that "the exact encoded output from Writer may be different from Go 1.26 as a result of the encoder implementation change", and that this carries through to &lt;code&gt;compress/gzip&lt;/code&gt;. The &lt;a href="https://github.com/opentofu/opentofu/blob/v1.13/CHANGELOG.md" rel="noopener noreferrer"&gt;OpenTofu 1.13 changelog&lt;/a&gt; calls the new output "equivalent to &lt;em&gt;but not equal to&lt;/em&gt;" the output of earlier releases. The new output is stable: three runs of 1.13.1 gave the same string.&lt;/p&gt;

&lt;p&gt;Terraform 1.16.5 is built with Go 1.26.8 and produced the same bytes as OpenTofu 1.12.7. So you also get this diff when you move from Terraform 1.16 to OpenTofu 1.13, not only when you upgrade OpenTofu.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the plan shows
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;cloud-init.yaml&lt;/strong&gt; unchanged&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;base64gzip()&lt;/strong&gt; Go 1.27 DEFLATE&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New string&lt;/strong&gt; same bytes after gunzip&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provider diff&lt;/strong&gt; string != state&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan&lt;/strong&gt; update or replace&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is the config we applied with 1.12.7:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;locals&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;user_data&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;base64gzip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"${path.module}/cloud-init.yaml"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Stands in for an attribute that a provider updates in place,&lt;/span&gt;
&lt;span class="c1"&gt;# such as user_data_base64 on aws_instance.&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"web_in_place"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;local&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;user_data&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Stands in for an attribute that forces replacement,&lt;/span&gt;
&lt;span class="c1"&gt;# such as custom_data on azurerm_linux_virtual_machine.&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"web_replace"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;triggers_replace&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;local&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;user_data&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After &lt;code&gt;tofu-1.12.7 apply&lt;/code&gt;, a 1.12.7 plan reported no changes. Then we planned with 1.13.1 against the same state and the same config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ tofu-1.13.1 plan
terraform_data.web_replace: Refreshing state... [id=4437d213-ef2e-bd8b-9c81-b5831d1b46c8]
terraform_data.web_in_place: Refreshing state... [id=b0ec5936-31c0-4653-0fb6-e353122675fa]

OpenTofu used the selected providers to generate the following execution
plan. Resource actions are indicated with the following symbols:
  ~ update in-place (current -&amp;gt; planned)
-/+ destroy and then create replacement

OpenTofu will perform the following actions:

  # terraform_data.web_in_place will be updated in-place
  ~ resource "terraform_data" "web_in_place" {
        id = "b0ec5936-31c0-4653-0fb6-e353122675fa"
      ~ input = "H4sIAAAAAAAA/5RU3U7zRhC9z..." -&amp;gt; "H4sIAAAAAAAA/5RTXW/jNhB89..."
      ~ output = "H4sIAAAAAAAA/5RU3U7zRhC9z..." -&amp;gt; (known after apply)
    }

  # terraform_data.web_replace must be replaced
-/+ resource "terraform_data" "web_replace" {
      ~ id = "4437d213-ef2e-bd8b-9c81-b5831d1b46c8" -&amp;gt; (known after apply)
      ~ triggers_replace = "H4sIAAAAAAAA/5RU3U7zRhC9z..." -&amp;gt; "H4sIAAAAAAAA/5RTXW/jNhB89..."
    }

Plan: 1 to add, 1 to change, 1 to destroy.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two resources, no edits, one update and one replacement. On real resources, the result depends on the provider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/instance" rel="noopener noreferrer"&gt;&lt;code&gt;aws_instance&lt;/code&gt;&lt;/a&gt;: for &lt;code&gt;user_data_base64&lt;/code&gt;, where gzip output belongs, "Updates to this field will trigger a stop/start of the EC2 instance by default. If the &lt;code&gt;user_data_replace_on_change&lt;/code&gt; is set then updates to this field will trigger a destroy and recreate of the EC2 instance."&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs/resources/linux_virtual_machine" rel="noopener noreferrer"&gt;&lt;code&gt;azurerm_linux_virtual_machine&lt;/code&gt;&lt;/a&gt;: for &lt;code&gt;custom_data&lt;/code&gt;, "Changing this forces a new resource to be created."&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/launch_template" rel="noopener noreferrer"&gt;&lt;code&gt;aws_launch_template&lt;/code&gt;&lt;/a&gt;: a changed &lt;code&gt;user_data&lt;/code&gt; creates a new template version. Instances get it the next time your Auto Scaling group launches or refreshes from that version.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a disaster if you plan for it. But a stop/start of a single production box at 2 p.m. is the kind of surprise you want to catch in plan review, not after someone approves an apply because "it is only a version bump".&lt;/p&gt;

&lt;h3&gt;
  
  
  Find it before the upgrade
&lt;/h3&gt;

&lt;p&gt;Start with a search. Run it after &lt;code&gt;tofu init&lt;/code&gt;, so that it also covers registry and git modules in &lt;code&gt;.terraform/modules&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rnE&lt;/span&gt; &lt;span class="nt"&gt;--include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'*.tf'&lt;/span&gt; &lt;span class="nt"&gt;--include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'*.tofu'&lt;/span&gt; &lt;span class="s1"&gt;'base64gzip\(|"winrm"'&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On our test folder it found both problems this post covers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tf&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;user_data&lt;/span&gt; &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;base64gzip&lt;/span&gt;&lt;span class="err"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="err"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"${path.module}/cloud-init.yaml"&lt;/span&gt;&lt;span class="err"&gt;))&lt;/span&gt;
&lt;span class="nx"&gt;winrm&lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tf&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"winrm"&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A search only finds literal calls. A plan is the reliable check. Run a 1.13.1 plan in a throwaway job, save it, and list each changed attribute with this &lt;code&gt;jq&lt;/code&gt; filter (save it as &lt;code&gt;changed-attrs.jq&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;.resource_changes[]
| &lt;span class="k"&gt;select&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;.change.actions &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"no-op"&lt;/span&gt;&lt;span class="o"&gt;])&lt;/span&gt;
| &lt;span class="nb"&gt;.&lt;/span&gt; as &lt;span class="nv"&gt;$r&lt;/span&gt;
| &lt;span class="o"&gt;[(&lt;/span&gt;&lt;span class="nv"&gt;$r&lt;/span&gt;.change.before // &lt;span class="o"&gt;{})&lt;/span&gt; | keys[]
    | &lt;span class="k"&gt;select&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$r&lt;/span&gt;.change.before[.] &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="nv"&gt;$r&lt;/span&gt;.change.after[.]
             and &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$r&lt;/span&gt;.change.after_unknown[.] | not&lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;
| &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="nv"&gt;$r&lt;/span&gt;&lt;span class="s2"&gt;.change.actions | join("&lt;/span&gt;+&lt;span class="s2"&gt;")) &lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="nv"&gt;$r&lt;/span&gt;&lt;span class="s2"&gt;.address) changed: &lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="s2"&gt;join("&lt;/span&gt;, &lt;span class="s2"&gt;"))"&lt;/span&gt;


&lt;span class="nv"&gt;$ &lt;/span&gt;tofu-1.13.1 plan &lt;span class="nt"&gt;-lock&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false&lt;/span&gt; &lt;span class="nt"&gt;-out&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;upgrade.tfplan &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null
&lt;span class="nv"&gt;$ &lt;/span&gt;tofu-1.13.1 show &lt;span class="nt"&gt;-json&lt;/span&gt; upgrade.tfplan | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; changed-attrs.jq
update terraform_data.web_in_place changed: input
delete+create terraform_data.web_replace changed: triggers_replace

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A plan does not write state, and &lt;code&gt;-lock=false&lt;/code&gt; stops a dry-run job from blocking a real apply. Remove the flag if you prefer to wait for the lock. You want a list in which every changed attribute is user data. Anything else in the list is real drift or a different upgrade change, so examine it separately.&lt;/p&gt;

&lt;h3&gt;
  
  
  Three ways to handle it
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Handling the base64gzip diff&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accept it&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Plan with 1.13.1, review, and apply in a window you choose.&lt;/span&gt;
&lt;span class="c"&gt;# The content is identical, so the only effect is the&lt;/span&gt;
&lt;span class="c"&gt;# stop/start or replacement itself.&lt;/span&gt;
tofu plan &lt;span class="nt"&gt;-out&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;upgrade.tfplan
tofu show &lt;span class="nt"&gt;-json&lt;/span&gt; upgrade.tfplan | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; changed-attrs.jq
tofu apply upgrade.tfplan

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Hide it (temporary)&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_instance"&lt;/span&gt; &lt;span class="s2"&gt;"web"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;# ...&lt;/span&gt;
  &lt;span class="nx"&gt;user_data_base64&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;base64gzip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"${path.module}/cloud-init.yaml"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

  &lt;span class="nx"&gt;lifecycle&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;# TEMPORARY: OpenTofu 1.13 re-encodes base64gzip output.&lt;/span&gt;
    &lt;span class="c1"&gt;# This also hides real cloud-init edits. Remove it on the next change.&lt;/span&gt;
    &lt;span class="nx"&gt;ignore_changes&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;user_data_base64&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Gzip in the provider&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="s2"&gt;"cloudinit_config"&lt;/span&gt; &lt;span class="s2"&gt;"web"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;gzip&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="nx"&gt;base64_encode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

  &lt;span class="nx"&gt;part&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;content_type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"text/cloud-config"&lt;/span&gt;
    &lt;span class="nx"&gt;content&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"${path.module}/cloud-init.yaml"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_instance"&lt;/span&gt; &lt;span class="s2"&gt;"web"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;# ...&lt;/span&gt;
  &lt;span class="nx"&gt;user_data_base64&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cloudinit_config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;web&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rendered&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Accept it.&lt;/strong&gt; This is usually the correct choice. Do the upgrade apply in a maintenance window, one environment at a time, and use the &lt;code&gt;jq&lt;/code&gt; list as the change record.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hide it.&lt;/strong&gt; The release notes suggest &lt;code&gt;ignore_changes&lt;/code&gt; as a temporary fix, and it works: with &lt;code&gt;ignore_changes&lt;/code&gt; on both test resources, the 1.13.1 plan said &lt;code&gt;No changes&lt;/code&gt;. Then we added a line to &lt;code&gt;cloud-init.yaml&lt;/code&gt; and planned again. It still said &lt;code&gt;No changes&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ignore_changes&lt;/code&gt; cannot tell the encoder change from a real edit. While it is in place, changes to your cloud-init file do not reach your instances, and the plan does not tell you. If you use it, open a ticket to remove it, and remove it on the next intentional user data change.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Move the gzip into the provider.&lt;/strong&gt; With &lt;code&gt;cloudinit_config&lt;/code&gt; and &lt;code&gt;gzip = true&lt;/code&gt;, the compression runs in the provider binary, not in OpenTofu. We rendered the same file with &lt;code&gt;hashicorp/cloudinit&lt;/code&gt; v2.4.1 under both 1.12.7 and 1.13.1, and the SHA256 of &lt;code&gt;rendered&lt;/code&gt; was identical. This does not make you immune. It moves the dependency: v2.4.1 is built with Go 1.26.8, and a future provider release built with Go 1.27 can cause the same one-time diff. So pin the provider version and read its changelog. The switch itself also changes your user data once, because &lt;code&gt;cloudinit_config&lt;/code&gt; wraps parts in a MIME multi-part document. Do it in the same window as the upgrade.&lt;/p&gt;

&lt;h2&gt;
  
  
  WinRM: passes validate and plan, fails at apply
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://devops-daily.com/posts/opentofu-1-12-destroy-false-state-surgery" rel="noopener noreferrer"&gt;OpenTofu 1.12&lt;/a&gt; deprecated the &lt;code&gt;winrm&lt;/code&gt; connection type, and 1.13 removed it (&lt;a href="https://github.com/opentofu/opentofu/pull/4012" rel="noopener noreferrer"&gt;#4012&lt;/a&gt;) because "some of the upstream libraries OpenTofu was using to implement these features are no longer maintained". We expected &lt;code&gt;tofu validate&lt;/code&gt; to report it. It does not. This is the test config (the host is a closed local port with a short timeout, so that the run fails fast):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"bootstrap"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;provisioner&lt;/span&gt; &lt;span class="s2"&gt;"remote-exec"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;inline&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"powershell -Command Install-WindowsFeature Web-Server"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="nx"&gt;connection&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"winrm"&lt;/span&gt;
      &lt;span class="nx"&gt;host&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"127.0.0.1"&lt;/span&gt;
      &lt;span class="nx"&gt;user&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"Administrator"&lt;/span&gt;
      &lt;span class="nx"&gt;password&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"example-only"&lt;/span&gt;
      &lt;span class="nx"&gt;https&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
      &lt;span class="nx"&gt;timeout&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"10s"&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;tofu-1&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;13.1&lt;/span&gt; &lt;span class="nx"&gt;validate&lt;/span&gt;
&lt;span class="nx"&gt;Warning&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;WinRM&lt;/span&gt; &lt;span class="nx"&gt;connection&lt;/span&gt; &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;is&lt;/span&gt; &lt;span class="nx"&gt;deprecated&lt;/span&gt;

  &lt;span class="nx"&gt;on&lt;/span&gt; &lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tf&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"bootstrap"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
   &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;connection&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
   &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"winrm"&lt;/span&gt;

&lt;span class="nx"&gt;The&lt;/span&gt; &lt;span class="nx"&gt;winrm&lt;/span&gt; &lt;span class="nx"&gt;connection&lt;/span&gt; &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;is&lt;/span&gt; &lt;span class="nx"&gt;deprecated&lt;/span&gt; &lt;span class="nx"&gt;and&lt;/span&gt; &lt;span class="nx"&gt;will&lt;/span&gt; &lt;span class="nx"&gt;be&lt;/span&gt; &lt;span class="nx"&gt;removed&lt;/span&gt; &lt;span class="nx"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;future&lt;/span&gt;
&lt;span class="nx"&gt;version&lt;/span&gt; &lt;span class="nx"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;OpenTofu&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;
&lt;span class="err"&gt;...&lt;/span&gt;
&lt;span class="nx"&gt;Success&lt;/span&gt;&lt;span class="err"&gt;!&lt;/span&gt; &lt;span class="nx"&gt;The&lt;/span&gt; &lt;span class="nx"&gt;configuration&lt;/span&gt; &lt;span class="nx"&gt;is&lt;/span&gt; &lt;span class="nx"&gt;valid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;but&lt;/span&gt; &lt;span class="nx"&gt;there&lt;/span&gt; &lt;span class="nx"&gt;were&lt;/span&gt; &lt;span class="nx"&gt;some&lt;/span&gt; &lt;span class="nx"&gt;validation&lt;/span&gt; &lt;span class="nx"&gt;warnings&lt;/span&gt;
&lt;span class="nx"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;shown&lt;/span&gt; &lt;span class="nx"&gt;above&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;

&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;tofu-1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;13.1&lt;/span&gt; &lt;span class="nx"&gt;plan&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nx"&gt;grep&lt;/span&gt; &lt;span class="s1"&gt;'Plan:'&lt;/span&gt;
&lt;span class="nx"&gt;Plan&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;add&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;change&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;destroy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;time&lt;/span&gt; &lt;span class="nx"&gt;tofu-1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;13.1&lt;/span&gt; &lt;span class="nx"&gt;apply&lt;/span&gt; &lt;span class="nx"&gt;-auto-approve&lt;/span&gt;
&lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="nx"&gt;Error&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;remote-exec&lt;/span&gt; &lt;span class="nx"&gt;provisioner&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;

  &lt;span class="nx"&gt;with&lt;/span&gt; &lt;span class="nx"&gt;terraform_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;bootstrap&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;on&lt;/span&gt; &lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tf&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"bootstrap"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
   &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;provisioner&lt;/span&gt; &lt;span class="s2"&gt;"remote-exec"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

&lt;span class="s1"&gt;'winrm'&lt;/span&gt; &lt;span class="nx"&gt;connections&lt;/span&gt; &lt;span class="nx"&gt;are&lt;/span&gt; &lt;span class="nx"&gt;not&lt;/span&gt; &lt;span class="nx"&gt;supported&lt;/span&gt; &lt;span class="nx"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;OpenTofu&lt;/span&gt; &lt;span class="nx"&gt;v1&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;13&lt;/span&gt; &lt;span class="nx"&gt;or&lt;/span&gt; &lt;span class="nx"&gt;later&lt;/span&gt;

&lt;span class="nx"&gt;Error&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Provisioners&lt;/span&gt; &lt;span class="nx"&gt;no&lt;/span&gt; &lt;span class="nx"&gt;longer&lt;/span&gt; &lt;span class="nx"&gt;support&lt;/span&gt; &lt;span class="nx"&gt;WinRM&lt;/span&gt;
&lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="nx"&gt;real&lt;/span&gt; &lt;span class="mi"&gt;0m&lt;/span&gt;&lt;span class="mf"&gt;0.153&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;

&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;tofu-1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;13.1&lt;/span&gt; &lt;span class="nx"&gt;show&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nx"&gt;head&lt;/span&gt; &lt;span class="nx"&gt;-2&lt;/span&gt;
&lt;span class="c1"&gt;# terraform_data.bootstrap: (tainted)&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"bootstrap"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;1.13.1 still prints the 1.12 deprecation warning at validate and plan. It says the feature "will be removed in a future version", but the feature is already gone in this version. The error comes in under a second at apply. For comparison, 1.12.7 tried to connect to port 5986 and failed at the 10-second timeout, as it should.&lt;/p&gt;

&lt;p&gt;The order matters. Provisioners run after the resource is created, and the &lt;a href="https://opentofu.org/docs/language/resources/provisioners/syntax/" rel="noopener noreferrer"&gt;OpenTofu docs&lt;/a&gt; say "if a creation-time provisioner fails, the resource is marked as &lt;strong&gt;tainted&lt;/strong&gt;" and "will be planned for destruction and recreation upon the next &lt;code&gt;tofu apply&lt;/code&gt;". With a real Windows VM, the VM is created and billed, then marked for replacement, and each apply after that recreates it and fails again until you remove the provisioner. Existing VMs whose provisioners ran long ago are not affected until something replaces them. Be careful with that last point: if a Windows VM uses gzipped &lt;code&gt;custom_data&lt;/code&gt;, the base64gzip change above is exactly the kind of thing that replaces it, and then its WinRM provisioner runs again and fails.&lt;/p&gt;

&lt;p&gt;To fix it, move to SSH or remove the provisioner:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Windows Server 2019 and later can run &lt;a href="https://learn.microsoft.com/en-us/windows-server/administration/openssh/openssh_install_firstuse" rel="noopener noreferrer"&gt;OpenSSH Server&lt;/a&gt;. Set &lt;code&gt;type = "ssh"&lt;/code&gt; and &lt;code&gt;target_platform = "windows"&lt;/code&gt; in the &lt;code&gt;connection&lt;/code&gt; block. If the SSH default shell is PowerShell, the &lt;a href="https://opentofu.org/docs/language/resources/provisioners/connection/" rel="noopener noreferrer"&gt;connection docs&lt;/a&gt; also tell you to set &lt;code&gt;script_path&lt;/code&gt; to a &lt;code&gt;.ps1&lt;/code&gt; path.&lt;/li&gt;
&lt;li&gt;Better, if you can: bake the configuration into the image, or run it from &lt;code&gt;custom_data&lt;/code&gt; or &lt;code&gt;user_data&lt;/code&gt;, so that no provisioner has to connect at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The grep above finds literal &lt;code&gt;"winrm"&lt;/code&gt; strings. It does not find &lt;code&gt;type = var.connection_type&lt;/code&gt;, so also search for &lt;code&gt;connection&lt;/code&gt; blocks whose type comes from a variable.&lt;/p&gt;

&lt;h2&gt;
  
  
  32-bit builds: last series
&lt;/h2&gt;

&lt;p&gt;The 1.13 changelog says this is "the final release series that will include official builds for 32-bit CPU architectures" (&lt;code&gt;*_386&lt;/code&gt; and &lt;code&gt;*_arm&lt;/code&gt;). The Pi's 64-bit kernel can run 32-bit ARM binaries, so we ran the &lt;code&gt;linux_arm&lt;/code&gt; build:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt;
&lt;span class="go"&gt;aarch64
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;tofu-1.13.1-linux-arm version
&lt;span class="go"&gt;OpenTofu v1.13.1
on linux_arm
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;tofu-1.13.1-linux-arm init
&lt;span class="go"&gt;
Warning: Support for 32-bit CPU architectures is ending soon

OpenTofu v1.13 is the last release series that will include official release
packages for 32-bit CPU architectures.

We recommend planning to migrate to a 64-bit CPU architecture instead.
Alternatively, you could build OpenTofu for linux_arm from source code
yourself, ...

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at the second line of &lt;code&gt;tofu version&lt;/code&gt; on each runner. If it says &lt;code&gt;on linux_arm&lt;/code&gt; or &lt;code&gt;on linux_386&lt;/code&gt;, that runner has to move. Common cases are 32-bit Raspberry Pi OS runners, old ARMv7 boards, and &lt;code&gt;i386&lt;/code&gt; container images. As the test shows, a 64-bit kernel does not help if the image or the binary you install is 32-bit. You have time: per the changelogs, the 1.13 series is supported until August 1, 2027, and 1.12 until February 1, 2027. The 1.11 series lost support on August 1, 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  The new functions: hints for unknown values
&lt;/h2&gt;

&lt;p&gt;This is the main feature of the release. During a plan, any value that an API assigns at creation time is unknown, and OpenTofu cannot use an unknown value to decide how many instances to create. Here is the classic case: a network module creates a VPC, and a second module creates flow logs only when it gets a VPC ID.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight terraform"&gt;&lt;code&gt;&lt;span class="c1"&gt;# modules/network/main.tf&lt;/span&gt;
&lt;span class="c1"&gt;# terraform_data stands in for aws_vpc: its output is unknown until apply,&lt;/span&gt;
&lt;span class="c1"&gt;# the same way a VPC id is decided by the AWS API at creation time.&lt;/span&gt;
&lt;span class="k"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"vpc"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"vpc-0a1b2c3d4e5f60718"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;output&lt;/span&gt; &lt;span class="s2"&gt;"vpc_id"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;terraform_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;output&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# modules/flow_logs/main.tf&lt;/span&gt;
&lt;span class="k"&gt;variable&lt;/span&gt; &lt;span class="s2"&gt;"vpc_id"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;string&lt;/span&gt;
  &lt;span class="nx"&gt;default&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"terraform_data"&lt;/span&gt; &lt;span class="s2"&gt;"flow_log"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;count&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kd"&gt;var&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpc_id&lt;/span&gt; &lt;span class="err"&gt;!&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="err"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
  &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kd"&gt;var&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpc_id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a first plan, 1.12.7 and 1.13.1 fail the same way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: Invalid count argument

  on modules/flow_logs/main.tf line 7, in resource "terraform_data" "flow_log":
   7: count = var.vpc_id != null ? 1 : 0

The "count" value depends on resource attributes that cannot be determined
until apply, so OpenTofu cannot predict how many instances will be created.
...

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://opentofu.org/docs/language/functions/assume_family/" rel="noopener noreferrer"&gt;&lt;code&gt;assume...&lt;/code&gt; functions&lt;/a&gt; let the module author state what is true about the value even before it exists. Our first attempt failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: Invalid function argument

  on modules/network/main.tf line 8, in output "vpc_id":
   8: value = assumenotnull(terraform_data.vpc.output)

Invalid value for "value" parameter: given value must have a known type;
consider using the \"convert\" function to specify a type to assume.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;terraform_data.output&lt;/code&gt; takes on the type of &lt;code&gt;input&lt;/code&gt;, so its type is also unknown during the plan. Most provider attributes, such as &lt;code&gt;aws_vpc.id&lt;/code&gt;, are typed strings and do not have this problem. For a value like this, the docs tell you to combine the hint with the new &lt;a href="https://opentofu.org/docs/language/functions/convert/" rel="noopener noreferrer"&gt;&lt;code&gt;convert&lt;/code&gt;&lt;/a&gt; function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="s2"&gt;"vpc_id"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;assumenotnull&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;terraform_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;string&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;tofu-1&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;13.1&lt;/span&gt; &lt;span class="nx"&gt;plan&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nx"&gt;grep&lt;/span&gt; &lt;span class="nx"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'will be created|Plan:'&lt;/span&gt;
  &lt;span class="c1"&gt;# module.flow_logs.terraform_data.flow_log[0] will be created&lt;/span&gt;
  &lt;span class="c1"&gt;# module.network.terraform_data.vpc will be created&lt;/span&gt;
&lt;span class="nx"&gt;Plan&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;add&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;change&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;destroy&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;

&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;tofu-1&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;13.1&lt;/span&gt; &lt;span class="nx"&gt;apply&lt;/span&gt; &lt;span class="nx"&gt;-auto-approve&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nx"&gt;tail&lt;/span&gt; &lt;span class="nx"&gt;-1&lt;/span&gt;
&lt;span class="nx"&gt;Apply&lt;/span&gt; &lt;span class="nx"&gt;complete&lt;/span&gt;&lt;span class="err"&gt;!&lt;/span&gt; &lt;span class="nx"&gt;Resources&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="nx"&gt;added&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="nx"&gt;changed&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="nx"&gt;destroyed&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The usual workaround for this error is a two-step apply with &lt;code&gt;-exclude&lt;/code&gt; (the error message itself suggests it). Here, one plan is enough.&lt;/p&gt;

&lt;p&gt;A hint is a promise, and OpenTofu checks it. The docs say that if the value turns out to be null, "the function raises an error", so the apply fails. Use a hint only where the provider really guarantees it, for example that an ID is never null after create.&lt;/p&gt;

&lt;h3&gt;
  
  
  Catch wiring mistakes at plan time
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;assumestringprefix&lt;/code&gt; is useful together with variable validation. Here the caller connects the subnet output to an input that expects a VPC ID, a mistake that is easy to miss in review:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight terraform"&gt;&lt;code&gt;&lt;span class="c1"&gt;# main.tf&lt;/span&gt;
&lt;span class="k"&gt;module&lt;/span&gt; &lt;span class="s2"&gt;"flow_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;source&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"./modules/flow_logs"&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;module&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;network&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;subnet_id&lt;/span&gt; &lt;span class="c1"&gt;# wrong output wired in&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# modules/flow_logs/main.tf&lt;/span&gt;
&lt;span class="k"&gt;variable&lt;/span&gt; &lt;span class="s2"&gt;"vpc_id"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;string&lt;/span&gt;

  &lt;span class="nx"&gt;validation&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;condition&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;var&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"vpc-"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nx"&gt;error_message&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"vpc_id must be a VPC id (vpc-...)."&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without hints, the plan passes because the value is unknown, and validation waits for apply. The apply created the VPC and the subnet, and then failed. Both outputs are filtered to the key lines, and the IDs are shortened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;tofu-1.13.1 plan
&lt;span class="go"&gt;Plan: 3 to add, 0 to change, 0 to destroy.
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;tofu-1.13.1 apply &lt;span class="nt"&gt;-auto-approve&lt;/span&gt;
&lt;span class="go"&gt;module.network.terraform_data.subnet: Creation complete after 0s [id=...]
module.network.terraform_data.vpc: Creation complete after 0s [id=...]
Error: Invalid value for variable
vpc_id must be a VPC id (vpc-...).
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;tofu-1.13.1 state list
&lt;span class="go"&gt;module.network.terraform_data.subnet
module.network.terraform_data.vpc

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then we added prefix hints to the network module outputs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="s2"&gt;"vpc_id"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;assumenotnull&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;assumestringprefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;terraform_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;string&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="s2"&gt;"vpc-"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="s2"&gt;"subnet_id"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;assumenotnull&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;assumestringprefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;terraform_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;string&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="s2"&gt;"subnet-"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the same mistake fails the plan, before anything is created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ tofu-1.13.1 plan
Error: Invalid value for variable

  on main.tf line 7, in module "flow_logs":
   7: vpc_id = module.network.subnet_id # wrong output wired in
    ├────────────────
    │ var.vpc_id is a string

vpc_id must be a VPC id (vpc-...).

This was checked by the validation rule at modules/flow_logs/main.tf:4,3-13.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you maintain shared modules, the outputs of your network, IAM and DNS modules are the best places for these hints. Callers get the benefit without changing their own code. &lt;code&gt;assumeequal&lt;/code&gt; goes further: the docs show it with the AWS provider's &lt;code&gt;arn_build&lt;/code&gt; function to make an IAM role ARN fully known at plan time, so that policy checks can see the real policy document.&lt;/p&gt;

&lt;h3&gt;
  
  
  Guard the module version correctly
&lt;/h3&gt;

&lt;p&gt;A module that uses these functions does not work on older versions, and the errors are not clear. On 1.12.7, &lt;code&gt;assumenotnull(...)&lt;/code&gt; alone gives &lt;code&gt;Call to unknown function&lt;/code&gt;. With &lt;code&gt;convert(..., string)&lt;/code&gt;, the error is &lt;code&gt;Invalid reference&lt;/code&gt;, because 1.12 reads &lt;code&gt;string&lt;/code&gt; as a resource address. Terraform 1.16.5 also gives &lt;code&gt;Call to unknown function&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;You would usually add &lt;code&gt;required_version = "&amp;gt;= 1.13.0"&lt;/code&gt; to a &lt;code&gt;terraform&lt;/code&gt; block. That does not work here. Since 1.12, OpenTofu ignores &lt;code&gt;required_version&lt;/code&gt; in &lt;code&gt;.tf&lt;/code&gt; files and honors it only in &lt;code&gt;.tofu&lt;/code&gt; files (see the &lt;a href="https://github.com/opentofu/opentofu/issues/3300" rel="noopener noreferrer"&gt;RFC tracking issue&lt;/a&gt; and the &lt;a href="https://opentofu.org/docs/language/settings/" rel="noopener noreferrer"&gt;settings docs&lt;/a&gt;). We tested this with a constraint that no version can meet:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Constraint and file&lt;/th&gt;
&lt;th&gt;1.12.7&lt;/th&gt;
&lt;th&gt;1.13.1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;required_version = "&amp;gt;= 99.0"&lt;/code&gt; in &lt;code&gt;main.tf&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;ignored, validate passes&lt;/td&gt;
&lt;td&gt;ignored, validate passes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;required_version = "&amp;gt;= 99.0"&lt;/code&gt; in &lt;code&gt;versions.tofu&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Incompatible module&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Incompatible module&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;language { compatible_with { opentofu = "&amp;gt;= 1.13" } }&lt;/code&gt; in &lt;code&gt;versions.tofu&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Incompatible module&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;plan passes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So put the guard in a &lt;code&gt;.tofu&lt;/code&gt; file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# versions.tofu (the language block needs OpenTofu 1.12 or later)&lt;/span&gt;
&lt;span class="nx"&gt;language&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;compatible_with&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;opentofu&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"&amp;gt;= 1.13"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On 1.12.7 this gives &lt;code&gt;This module is not compatible with OpenTofu v1.12.7&lt;/code&gt;, which is the clear message you want. If a module must also work with Terraform, use the &lt;code&gt;.tofu&lt;/code&gt; precedence rule: when &lt;code&gt;outputs.tf&lt;/code&gt; and &lt;code&gt;outputs.tofu&lt;/code&gt; are both present, OpenTofu loads only the &lt;code&gt;.tofu&lt;/code&gt; file. We put a hinted output in &lt;code&gt;outputs.tofu&lt;/code&gt; and a plain one in &lt;code&gt;outputs.tf&lt;/code&gt;. &lt;code&gt;tofu validate&lt;/code&gt; (1.13.1) and &lt;code&gt;terraform validate&lt;/code&gt; (1.16.5) both passed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lint experiment
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;validate&lt;/code&gt;, &lt;code&gt;plan&lt;/code&gt;, &lt;code&gt;apply&lt;/code&gt; and &lt;code&gt;refresh&lt;/code&gt; accept a new &lt;code&gt;-lint&lt;/code&gt; flag. 1.13 has four rules, all for the root module only: &lt;code&gt;core:no-type-variable&lt;/code&gt;, &lt;code&gt;core:unused-variable&lt;/code&gt;, &lt;code&gt;core:unused-local&lt;/code&gt;, and &lt;code&gt;core:count-instead-enabled&lt;/code&gt;. The last one suggests the &lt;code&gt;enabled&lt;/code&gt; lifecycle argument instead of &lt;code&gt;count = cond ? 1 : 0&lt;/code&gt;. We ran the experiment on a small file that breaks all four rules:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;tofu-1.13.1 validate &lt;span class="nt"&gt;-lint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;all &lt;span class="nt"&gt;-json&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.diagnostics[] | select(.severity == "warning") | .summary'&lt;/span&gt;
Experimental linting enabled
Input variable not used &lt;span class="o"&gt;(&lt;/span&gt;core:unused-variable&lt;span class="o"&gt;)&lt;/span&gt;
Local value not used &lt;span class="o"&gt;(&lt;/span&gt;core:unused-local&lt;span class="o"&gt;)&lt;/span&gt;
Could use enabled instead of count &lt;span class="o"&gt;(&lt;/span&gt;core:count-instead-enabled&lt;span class="o"&gt;)&lt;/span&gt;
Variable with no &lt;span class="nb"&gt;type&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;core:no-type-variable&lt;span class="o"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exit code was 0. Lint results are warnings, and OpenTofu always adds an "Experimental linting enabled" warning. You can turn off one rule with &lt;code&gt;!&lt;/code&gt;, for example &lt;code&gt;-lint='all,!core:unused-variable'&lt;/code&gt;. The &lt;a href="https://opentofu.org/docs/language/linting/" rel="noopener noreferrer"&gt;linting docs&lt;/a&gt; say that rules can change even in minor releases. For now, make it a non-blocking CI step and read the output. Do not gate merges on it yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Smaller changes worth a line
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Plan text.&lt;/strong&gt; After &lt;code&gt;No changes. Your infrastructure matches the configuration.&lt;/code&gt;, 1.12.7 printed one more paragraph ("OpenTofu has compared your real infrastructure ... no changes are needed."). 1.13.1 does not. If a script greps for that paragraph, change it to use &lt;code&gt;tofu plan -detailed-exitcode&lt;/code&gt;, which returns 0 for no changes, 1 for errors, and 2 for changes. Our upgrade plan returned 2.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Platforms.&lt;/strong&gt; Windows on ARM64 is now an official platform, and macOS builds require macOS 13 Ventura or later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State encryption.&lt;/strong&gt; The &lt;code&gt;aws_kms&lt;/code&gt; key provider accepts an &lt;code&gt;encryption_context&lt;/code&gt;. &lt;code&gt;gcp_kms&lt;/code&gt; accepts &lt;code&gt;additional_authenticated_data&lt;/code&gt;, and &lt;code&gt;openbao&lt;/code&gt; accepts &lt;code&gt;associated_data&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crash recovery.&lt;/strong&gt; A Go panic now writes a partial &lt;code&gt;errored.tfstate&lt;/code&gt; to help you recover.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Saved plans&lt;/strong&gt; include the provider schemas, so &lt;code&gt;tofu show&lt;/code&gt; on a plan file usually does not need to start providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you stay on 1.12 for now&lt;/strong&gt; , take &lt;a href="https://github.com/opentofu/opentofu/releases/tag/v1.12.7" rel="noopener noreferrer"&gt;1.12.7&lt;/a&gt;. It fixes a deadlock that an attacker-controlled SSH server could cause in &lt;code&gt;remote-exec&lt;/code&gt; and &lt;code&gt;file&lt;/code&gt; provisioners (CVE-2026-78662).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Upgrade checklist
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;After &lt;code&gt;tofu init&lt;/code&gt;, search your code and &lt;code&gt;.terraform/modules&lt;/code&gt; for &lt;code&gt;base64gzip(&lt;/code&gt; and &lt;code&gt;"winrm"&lt;/code&gt;, and look for &lt;code&gt;connection&lt;/code&gt; blocks whose type comes from a variable.&lt;/li&gt;
&lt;li&gt;Remove WinRM provisioners &lt;strong&gt;before&lt;/strong&gt; you upgrade. Validate and plan will not stop you, and a failure at apply taints the resource.&lt;/li&gt;
&lt;li&gt;Run a 1.13.1 plan against real state in a throwaway job, save it, and run the &lt;code&gt;jq&lt;/code&gt; filter. Every changed attribute should be user data.&lt;/li&gt;
&lt;li&gt;For each user data change, decide: accept the stop/start or replacement in a window, or use a temporary &lt;code&gt;ignore_changes&lt;/code&gt; with a ticket to remove it.&lt;/li&gt;
&lt;li&gt;Check &lt;code&gt;tofu version&lt;/code&gt; on every runner and workstation image, and move anything on &lt;code&gt;linux_arm&lt;/code&gt; or &lt;code&gt;linux_386&lt;/code&gt; to 64-bit before August 1, 2027.&lt;/li&gt;
&lt;li&gt;Change scripts that read plan text to use &lt;code&gt;-detailed-exitcode&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;In shared modules, add &lt;code&gt;assume...&lt;/code&gt; hints to outputs whose IDs callers use in &lt;code&gt;count&lt;/code&gt;, &lt;code&gt;for_each&lt;/code&gt; or validation. Put a &lt;code&gt;language&lt;/code&gt; block in a &lt;code&gt;.tofu&lt;/code&gt; file to guard them.&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;-lint=all&lt;/code&gt; to CI as a non-blocking step and see what it finds.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The upgrade is not hard, but two of these changes are not visible until a specific step: the base64gzip diff shows only in a plan against real state, and the WinRM removal shows only at apply. If you run steps 1 to 3 first, you see both before they affect a real server.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/opentofu-1-13-upgrade-base64gzip-winrm" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>terraform</category>
      <category>opentofu</category>
      <category>infrastructureascode</category>
      <category>cloudinit</category>
    </item>
    <item>
      <title>p99, Load Balancers and Autoscaling: Latency Intuition You Can Play With</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Tue, 06 Oct 2026 18:43:17 +0000</pubDate>
      <link>https://dev.to/devopsdaily/p99-load-balancers-and-autoscaling-latency-intuition-you-can-play-with-41mk</link>
      <guid>https://dev.to/devopsdaily/p99-load-balancers-and-autoscaling-latency-intuition-you-can-play-with-41mk</guid>
      <description>&lt;p&gt;Averages hide latency problems. Every SRE learns this, usually the hard way: the dashboard says 80ms average while support tickets say the app is slow, because the average smooths over a tail where one request in a hundred takes four seconds, and your heaviest users, the ones making the most requests, hit that tail most often.&lt;/p&gt;

&lt;p&gt;Percentiles are how you see the tail, and the vocabulary around them (p50, p95, p99, tail latency) is quick to memorize and slow to internalize. Intuition, the ability to look at a p99 spike and have sensible suspects, usually comes from incidents. Three free browser simulators let you build a chunk of it without the incidents. Disclosure: I help build them; free, browser-based, no signup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part 1: watch a percentile move
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://devops-daily.com/games/latency-percentiles-simulator" rel="noopener noreferrer"&gt;Latency Percentiles Simulator&lt;/a&gt; generates request-latency distributions and renders them as a histogram with P50, P90, P95 and P99 markers that move as you change the controls. The five scenarios are synthetic distributions shaped like five kinds of trouble:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Healthy API&lt;/strong&gt;: the baseline. A tight distribution; the percentiles sit close together. This is what "fine" looks like, and knowing its shape is what lets you recognize not-fine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cold Starts&lt;/strong&gt;: a small fraction of requests pay a big startup cost. P50 barely notices; P99 jumps. This is the canonical "the average is fine, the tail is not" case, and it is the shape serverless platforms and JIT warmup produce.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache Churn&lt;/strong&gt;: hits are fast, misses are slow, and the mix moves the middle percentiles around. Different shape than cold starts: the distribution goes wide rather than growing a spike.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Noisy Neighbor&lt;/strong&gt;: shared infrastructure steals capacity intermittently, smearing the whole distribution outward. The tail grows, but unlike cold starts, so does everything above the median.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rolling Deploy&lt;/strong&gt;: a distribution shaped like a fleet mid-deploy, two populations blended, which is the shape to recognize in real telemetry when half your instances are new.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These are stylized shapes, not simulations of caches or neighbors, and that is fine for the purpose: the transferable skill is that similar p99 numbers can sit on very different distributions. Percentiles alone do not diagnose a cause, but knowing the classic shapes gives you better suspects to check when a tail spike arrives.&lt;/p&gt;

&lt;p&gt;The other lesson hiding here is why tails matter more than their percentage suggests: if a page fans out to dozens of roughly independent backend calls and waits for all of them, the chance that at least one hits the p99 grows fast, so the "1% case" ends up in far more than 1% of page loads. Google's SRE material makes the same point: a backend's p99 can effectively become the frontend's median.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part 2: the balancer decides who eats the tail
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://devops-daily.com/games/load-balancer-simulator" rel="noopener noreferrer"&gt;Load Balancer Simulator&lt;/a&gt; animates a three-server fleet behind a balancer you configure live: the algorithm (&lt;strong&gt;round robin, least connections, IP hash, random&lt;/strong&gt;), the traffic rate up to a burst mode, a server crash probability (up to a brutal 30%), retries on or off, and, the good part, you can click a server to knock it offline mid-run.&lt;/p&gt;

&lt;p&gt;Two experiments teach the most. First, crank the failure rate with retries on and watch what a crash actually costs: the request does not just fail, it goes around again, adding load precisely when the fleet is struggling. Retries are not free; the counter of crashed-then-retried requests makes that concrete. Second, knock a server offline under each algorithm and watch the traffic redistribute; then try it with &lt;strong&gt;IP hash&lt;/strong&gt;, where the same clients always land on the same server. Stickiness is what session-state architectures want, and the redistribution behavior when a sticky target dies is the price nobody mentions in the algorithm's one-line description.&lt;/p&gt;

&lt;p&gt;Round robin versus &lt;strong&gt;least connections&lt;/strong&gt; rounds it out: equal turns versus routing by current load. With identical healthy servers they look similar; under crashes and bursts, watching which algorithm piles work where is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part 3: capacity as a moving target
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://devops-daily.com/games/scaling-simulator" rel="noopener noreferrer"&gt;Scaling Simulator&lt;/a&gt; hands you autoscaling controls (thresholds, instance counts, vertical versus horizontal choices) and four traffic scenarios to survive: &lt;strong&gt;Gradual Growth, Sudden Spike, Black Friday, Variable Load&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Gradual growth is the tutorial level; almost any policy works. The hard lessons are in the other three:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sudden Spike&lt;/strong&gt; punishes reactive scaling: by the time the threshold trips and instances come up, the spike already hurt users. The lesson is that autoscaling has a reaction time, and traffic that moves faster than it needs headroom, not thresholds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Black Friday&lt;/strong&gt; is the predictable-surge case: adding capacity before the surge you know is coming beats any reactive policy. The simulator lets you provision ahead by hand, which is the whole trick, minus the calendar automation real platforms add.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Variable Load&lt;/strong&gt; is oscillating traffic, the pattern that punishes twitchy autoscaling with flapping in real systems. The simulator's scale-down cooldown is deliberately long, so what you actually observe is the defense working: capacity ratchets up and holds through the dips instead of thrashing. Cooldowns and stabilization windows are exactly how production autoscalers (Kubernetes' HPA included) buy calm at the cost of some idle capacity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The through-line with parts 1 and 2: capacity decisions surface as latency and errors. The scaling simulator tracks response time and failures rather than full percentile distributions, but the connection stands: under-provisioned fleets are where the ugly histograms of part 1 come from, and the autoscaler's job is to buy the healthy shape with as few idle instances as possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using the trio
&lt;/h2&gt;

&lt;p&gt;A sequence that works: learn the five distribution shapes in part 1, run the crash-and-retry and kill-a-server experiments in part 2, then survive Sudden Spike and Black Friday in part 3. After that, three things in production read differently: a p99 chart makes you ask what the distribution under it looks like, balancer choice reads as a failure-behavior decision, and autoscaler settings read as a bet about how your traffic moves.&lt;/p&gt;

&lt;p&gt;All three live with &lt;a href="https://devops-daily.com/games" rel="noopener noreferrer"&gt;the rest of our free DevOps games and simulators&lt;/a&gt;. For the deep end afterwards, Google's SRE book chapter on &lt;a href="https://sre.google/sre-book/monitoring-distributed-systems/" rel="noopener noreferrer"&gt;monitoring distributed systems&lt;/a&gt; and Gil Tene's classic &lt;a href="https://www.youtube.com/watch?v=lJ8ydIuPFeU" rel="noopener noreferrer"&gt;How NOT to Measure Latency&lt;/a&gt; are the standard references, easier to absorb once you have watched a histogram misbehave yourself.&lt;/p&gt;

</description>
      <category>sre</category>
      <category>performance</category>
      <category>devops</category>
      <category>architecture</category>
    </item>
    <item>
      <title>A Poisoned .git/config Runs Code on git status. We Tested Which Commands and Copies Carry It</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Tue, 06 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/a-poisoned-gitconfig-runs-code-on-git-status-we-tested-which-commands-and-copies-carry-it-2hok</link>
      <guid>https://dev.to/devopsdaily/a-poisoned-gitconfig-runs-code-on-git-status-we-tested-which-commands-and-copies-carry-it-2hok</guid>
      <description>&lt;p&gt;On October 2, GitLab's Threat Research Group published &lt;a href="https://about.gitlab.com/blog/deepseek-reasonix-vulnerability-discovered/" rel="noopener noreferrer"&gt;ConfigPoisoning&lt;/a&gt; (CVE-2026-102437), a command execution bug in DeepSeek-Reasonix Studio, a git client built for working next to AI coding assistants. The tool was careful. It neutralized &lt;code&gt;core.fsmonitor&lt;/code&gt; on every git call and added &lt;code&gt;--no-ext-diff --no-textconv&lt;/code&gt; to every diff. It still ran attacker code when you opened a diff, because a clean filter named in &lt;code&gt;.gitattributes&lt;/code&gt; runs anyway. GitLab says "the same pattern is present in several other widely used agent tools, currently under coordinated disclosure."&lt;/p&gt;

&lt;p&gt;It is the second disclosure of this kind in five weeks. On September 1, Manifold Security published &lt;a href="https://www.manifold.security/blog/ai-coding-agents-git-hijack" rel="noopener noreferrer"&gt;GitSpawn&lt;/a&gt;: eight findings across seven coding agents (Claude Code, Codex, Cursor, goose, Qwen Code, Grok Build and Hermes Agent) where repository-controlled git settings ran commands outside the agent's sandbox. Most of the detailed cases used a &lt;code&gt;core.fsmonitor&lt;/code&gt; line in the repo's own &lt;code&gt;.git/config&lt;/code&gt;; in Claude Code's case, the start-up &lt;code&gt;git status&lt;/code&gt; ran it while the workspace-trust prompt was still waiting for an answer.&lt;/p&gt;

&lt;p&gt;Both reports are about agents, but the mechanism is plain git, and git also runs in your CI jobs, bots and editor plugins. So we tested git itself: which everyday commands run programs that a repo's &lt;code&gt;.git/config&lt;/code&gt; or hooks point at, whether the usual hardening flags stop them, and which ways of moving a repo between machines carry that config along.&lt;/p&gt;

&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A bare &lt;code&gt;git status&lt;/code&gt; ran repo-supplied programs: &lt;code&gt;core.fsmonitor&lt;/code&gt;, a &lt;code&gt;post-index-change&lt;/code&gt; hook from &lt;code&gt;.git/hooks&lt;/code&gt; or &lt;code&gt;core.hooksPath&lt;/code&gt;, and on git 2.55 a hook defined in config.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git diff --no-ext-diff --no-textconv&lt;/code&gt; still ran the clean filter (and a &lt;code&gt;filter.&amp;lt;name&amp;gt;.process&lt;/code&gt; helper). Overrides only stopped a filter when we named its driver.&lt;/li&gt;
&lt;li&gt;On git 2.55, &lt;code&gt;-c core.hooksPath=/dev/null&lt;/code&gt; did not stop a hook defined in config (&lt;code&gt;hook.&amp;lt;name&amp;gt;.command&lt;/code&gt;, new in git 2.54). Adding &lt;code&gt;-c hook.&amp;lt;event&amp;gt;.enabled=false&lt;/code&gt; did, in our fixture.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git clone&lt;/code&gt; and cloning from a bundle never brought the config or hooks. A &lt;code&gt;tar&lt;/code&gt; of the working copy, which is what a workspace cache or artifact is, always did.&lt;/li&gt;
&lt;li&gt;git's ownership check (&lt;code&gt;safe.directory&lt;/code&gt;) refused a copy owned by another user on our Raspberry Pi. On GitHub-hosted runners it never fired, because the images set &lt;code&gt;safe.directory = *&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The scripts, the recorded results and a read-only audit are in the repo below.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;git and bash, to run the scripts&lt;/li&gt;
&lt;li&gt;A machine where you can create throwaway repos: every "payload" in the tests only appends a line to a log file&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why a repo can make git run programs at all
&lt;/h2&gt;

&lt;p&gt;Some git settings are commands, not values. &lt;code&gt;core.fsmonitor&lt;/code&gt; names a helper that tells git which files changed, so big repos do not stat every file. A filter driver (&lt;code&gt;filter.&amp;lt;name&amp;gt;.clean&lt;/code&gt;, &lt;code&gt;smudge&lt;/code&gt; or the long-running &lt;code&gt;process&lt;/code&gt;) rewrites file content on the way into and out of the index; that is how Git LFS works. &lt;code&gt;diff.external&lt;/code&gt; and &lt;code&gt;diff.&amp;lt;name&amp;gt;.textconv&lt;/code&gt; replace or preprocess diffs. Hooks are executables in &lt;code&gt;.git/hooks&lt;/code&gt; or in the directory &lt;code&gt;core.hooksPath&lt;/code&gt; points to, and since &lt;a href="https://github.com/git/git/blob/master/Documentation/RelNotes/2.54.0.adoc" rel="noopener noreferrer"&gt;git 2.54&lt;/a&gt; a hook can also be a command defined in config (&lt;code&gt;hook.&amp;lt;name&amp;gt;.command&lt;/code&gt; plus &lt;code&gt;hook.&amp;lt;name&amp;gt;.event&lt;/code&gt;, see &lt;a href="https://git-scm.com/docs/git-hook" rel="noopener noreferrer"&gt;git-hook&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Git reads all of these from the repository's own &lt;code&gt;.git/config&lt;/code&gt; and hooks directory. That is documented, intended behaviour for a repo you created. The trouble starts when &lt;code&gt;.git&lt;/code&gt; came from someone else.&lt;/p&gt;

&lt;p&gt;A normal clone protects you: git does not transfer &lt;code&gt;.git/config&lt;/code&gt; or hooks over a fetch. &lt;code&gt;.gitattributes&lt;/code&gt; does travel, because it is a tracked file, but a filter named there does nothing unless config defines it. So the question for a DevOps team is not "can a repo do this", it is "where do we copy a &lt;code&gt;.git&lt;/code&gt; directory instead of cloning it". GitLab's report lists the answers: "an archive, a synced folder, a CI cache, or a devcontainer build", plus a compromised agent that writes the config into a repo you cloned normally.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;The repo builds small throwaway repositories. Each has one committed file and the same file with a line appended, so there is always a change to show, and a &lt;code&gt;.git/config&lt;/code&gt; (or hooks directory) that points one setting at a script whose only job is to log that it ran and pass content through:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/The-DevOps-Daily/git-config-exec-check" rel="noopener noreferrer"&gt;The-DevOps-Daily/git-config-exec-check on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We tested eleven settings: &lt;code&gt;core.fsmonitor&lt;/code&gt;; clean, smudge and process filters; a &lt;code&gt;textconv&lt;/code&gt; driver; &lt;code&gt;diff.external&lt;/code&gt;; hooks in &lt;code&gt;.git/hooks&lt;/code&gt;; hooks via &lt;code&gt;core.hooksPath&lt;/code&gt;; a hook defined in config; &lt;code&gt;core.sshCommand&lt;/code&gt;; and &lt;code&gt;core.pager&lt;/code&gt;. For each one, &lt;code&gt;scripts/matrix.sh&lt;/code&gt; copies the repo fresh, runs 15 commands that tools and people run all the time, and records the exit code and what ran.&lt;/p&gt;

&lt;p&gt;We ran it on a Raspberry Pi with git 2.39.5, and on GitHub-hosted &lt;code&gt;ubuntu-latest&lt;/code&gt; and &lt;code&gt;macos-latest&lt;/code&gt; runners with git 2.55.0. Ubuntu and macOS were identical. Git 2.39.5 ran the same programs except for the config-hook column, which it does not support. This is the git 2.55.0 table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;fsmonitor&lt;/th&gt;
&lt;th&gt;clean&lt;/th&gt;
&lt;th&gt;process&lt;/th&gt;
&lt;th&gt;smudge&lt;/th&gt;
&lt;th&gt;textconv&lt;/th&gt;
&lt;th&gt;diff.external&lt;/th&gt;
&lt;th&gt;hooks dir&lt;/th&gt;
&lt;th&gt;config hook&lt;/th&gt;
&lt;th&gt;sshCommand&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git status&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;post-index-change&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git status --porcelain&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;post-index-change&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git diff&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git diff --no-ext-diff --no-textconv&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git log -p -1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git log --oneline -1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git show HEAD&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git blame notes.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git ls-files&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git rev-parse HEAD&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git add -A&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;post-index-change&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git commit -qam wip&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;post-commit, post-index-change, pre-commit, reference-transaction&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git stash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;post-index-change, reference-transaction&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git checkout -- notes.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;post-checkout, post-index-change&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git fetch origin (fails)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;How to read it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A blank cell means "not observed with this fixture", not "can never run". Because our modified file is longer than the committed one, git can see the change from the file size without filtering it. With a same-size edit or a touched file, &lt;code&gt;git status&lt;/code&gt; may need to compare content, and then filters can run too.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;hooks dir&lt;/code&gt; was identical for &lt;code&gt;.git/hooks&lt;/code&gt; and &lt;code&gt;core.hooksPath&lt;/code&gt;, so they share a column. &lt;code&gt;core.pager&lt;/code&gt; never ran, because none of the commands had a terminal; agents and CI jobs usually do not either.&lt;/li&gt;
&lt;li&gt;Our &lt;code&gt;process&lt;/code&gt; helper does not speak git's filter protocol. Git started it either way; git 2.39.5 then exited with 128, and git 2.55.0 carried on and exited 0.&lt;/li&gt;
&lt;li&gt;Only the &lt;code&gt;sshCommand&lt;/code&gt; fixture has a remote, and its fetch fails on purpose. A successful fetch can fire more hooks than this row shows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three rows deserve a second look.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;git status&lt;/code&gt; is not read-only.&lt;/strong&gt; It refreshes the index, and refreshing the index is when git asks the fsmonitor helper what changed. When the refresh writes the index back, git fires the &lt;code&gt;post-index-change&lt;/code&gt; hook, from the hooks directory and from config. All of that happened on a repo with one modified file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;git fetch&lt;/code&gt; ran fsmonitor before it contacted the remote.&lt;/strong&gt; In the fsmonitor fixture there is no remote at all. &lt;code&gt;scripts/trace-fetch.sh&lt;/code&gt; shows git starting the helper twice (asking for protocol version 2, then falling back to version 1 after our helper exited non-zero) and only then failing to find &lt;code&gt;origin&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git 2.39.5 on Linux aarch64
run-command.c:655 trace: run_command: cd &amp;lt;tmp&amp;gt;/r; '&amp;lt;tmp&amp;gt;/hook.sh fsmonitor' 2 1791269076974045040
run-command.c:655 trace: run_command: cd &amp;lt;tmp&amp;gt;/r; '&amp;lt;tmp&amp;gt;/hook.sh fsmonitor' 1 1791269076974045040
run-command.c:655 trace: run_command: unset GIT_PREFIX; GIT_PROTOCOL=version=2 'git-upload-pack '\''origin'\'''
fatal: 'origin' does not appear to be a git repository
fatal: Could not read from remote repository.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The diff flags that hardened tools add are aimed at the wrong half.&lt;/strong&gt; &lt;code&gt;--no-ext-diff&lt;/code&gt; and &lt;code&gt;--no-textconv&lt;/code&gt; stop the programs that render a diff. The clean filter runs earlier, when git turns the working-tree file into a blob to compare. That is the GitLab finding, and plain git 2.55 behaves the same way. Diffs between two commits (&lt;code&gt;git log -p&lt;/code&gt;, &lt;code&gt;git show&lt;/code&gt;) never ran a filter, because no working-tree file is involved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do the usual flags help?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;scripts/overrides.sh&lt;/code&gt; builds one repo with fsmonitor, a hook in &lt;code&gt;.git/hooks&lt;/code&gt;, a hook defined in config, &lt;code&gt;diff.external&lt;/code&gt;, a &lt;code&gt;textconv&lt;/code&gt; driver and a clean filter. Each row starts from a fresh copy, runs &lt;code&gt;git status&lt;/code&gt; and &lt;code&gt;git diff&lt;/code&gt;, and adds more protection. On git 2.55.0:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git 2.55.0 on Linux x86_64
plain exit 0/0 ran: diff-external,filter-clean,fsmonitor,hook:config,hook:dir
diff --no-ext-diff --no-textconv exit 0/0 ran: filter-clean,fsmonitor,hook:config,hook:dir
+ -c core.fsmonitor=false -c core.hooksPath=/dev/null exit 0/0 ran: filter-clean,hook:config
+ -c filter.lab.clean= (needs the driver name) exit 0/0 ran: hook:config
+ -c hook.post-index-change.enabled=false (per event) exit 0/0 ran:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read it row by row:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;With no protection, five programs ran: fsmonitor, the hooks-directory hook, the config hook, the clean filter and &lt;code&gt;diff.external&lt;/code&gt;. The &lt;code&gt;textconv&lt;/code&gt; driver did not, because &lt;code&gt;diff.external&lt;/code&gt; replaces the whole diff; the matrix above is where &lt;code&gt;textconv&lt;/code&gt; shows up, and where &lt;code&gt;--no-textconv&lt;/code&gt; stops it.&lt;/li&gt;
&lt;li&gt;The diff flags stopped &lt;code&gt;diff.external&lt;/code&gt;, and nothing else.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;core.fsmonitor=false&lt;/code&gt; and &lt;code&gt;core.hooksPath=/dev/null&lt;/code&gt; are what a careful wrapper adds. They stopped fsmonitor and the hooks-directory hook, but not the hook defined in config, which is not looked up in a hooks directory. The clean filter also still ran.&lt;/li&gt;
&lt;li&gt;Clearing the clean filter worked only because we knew the driver was called &lt;code&gt;lab&lt;/code&gt;. An attacker picks the name, and &lt;code&gt;.gitattributes&lt;/code&gt; can name a different driver for each file pattern. The same goes for &lt;code&gt;smudge&lt;/code&gt; and &lt;code&gt;process&lt;/code&gt;, which a wrapper has to clear too.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;hook.&amp;lt;event&amp;gt;.enabled=false&lt;/code&gt;, which the git 2.55 &lt;a href="https://git-scm.com/docs/git-hook" rel="noopener noreferrer"&gt;git-hook documentation&lt;/a&gt; describes as a switch for every hook of one event, stopped the config hook in this fixture without knowing its name. We only tested it together with &lt;code&gt;core.hooksPath=/dev/null&lt;/code&gt;, so keep both. It is also per event, so a wrapper has to list every event its commands can fire.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On git 2.39.5 the same script gave the same first four rows, minus the config hook, which that version does not know about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which copies bring the config along
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;scripts/delivery.sh&lt;/code&gt; builds one repo with &lt;code&gt;core.fsmonitor&lt;/code&gt;, a clean filter and a &lt;code&gt;post-index-change&lt;/code&gt; hook. All three call a helper stored inside &lt;code&gt;.git&lt;/code&gt; by relative path, so the payload travels with any copy that includes &lt;code&gt;.git&lt;/code&gt;. It copies the repo four ways, records anything that ran during the copy, and then runs &lt;code&gt;git status&lt;/code&gt; and &lt;code&gt;git diff&lt;/code&gt; in each copy. On the Pi:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git 2.39.5 on Linux aarch64
safe.directory already set on this machine: no
original repo copy ran: - then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0
git clone copy ran: nothing then ran: nothing exit 0/0
git clone from a bundle copy ran: nothing then ran: nothing exit 0/0
tar of the working copy copy ran: nothing then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0
same tar, owned by another user copy ran: - then ran: nothing exit 128/129 fatal: detected dubious ownership in repository at '/tmp/tmp.S42grY0m9P/other-owner'
  ...with no system or global config copy ran: - then ran: nothing exit 128/129 fatal: detected dubious ownership in repository at '/tmp/tmp.S42grY0m9P/other-owner'
  ...with -c safe.directory=* copy ran: - then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The clone and the bundle were clean, both during the copy and afterwards. The tar was not, and a tar of the working copy is what most "cache the workspace" setups produce: an &lt;code&gt;actions/cache&lt;/code&gt; or GitLab &lt;code&gt;cache:&lt;/code&gt; entry that includes &lt;code&gt;.git&lt;/code&gt;, a build artifact someone zipped from the job directory, a devcontainer volume, a folder synced between machines. Once that copy lands, the next tool to run &lt;code&gt;git status&lt;/code&gt; in it runs the repo's programs.&lt;/p&gt;

&lt;p&gt;Here is that as a terminal session, from &lt;code&gt;scripts/demo-restored-workspace.sh&lt;/code&gt;: a workspace restored from a tarball, two git commands, and our audit script at the end. The helper sits inside &lt;code&gt;.git&lt;/code&gt;, so it arrived with the tarball, and the script empties &lt;code&gt;ran.log&lt;/code&gt; before each git command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;restored workspace, git 2.39.5&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# a workspace restored from a cache tarball, not cloned&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git status &lt;span class="nt"&gt;--short&lt;/span&gt;
 M notes.txt
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; ../ran.log
ran: fsmonitor
ran: fsmonitor
&lt;span class="nv"&gt;$ &lt;/span&gt;git diff &lt;span class="nt"&gt;--no-ext-diff&lt;/span&gt; &lt;span class="nt"&gt;--no-textconv&lt;/span&gt; &lt;span class="nt"&gt;--stat&lt;/span&gt;
 notes.txt | 1 +
 1 file changed, 1 insertion&lt;span class="o"&gt;(&lt;/span&gt;+&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; ../ran.log
ran: fsmonitor
ran: fsmonitor
ran: filter-clean
ran: filter-clean
&lt;span class="nv"&gt;$ &lt;/span&gt;audit.sh .&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"exit &lt;/span&gt;&lt;span class="nv"&gt;$?&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
Repo config &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; can run programs:
file:.git/config core.fsmonitor .git/lab-hook.sh fsmonitor
file:.git/config filter.lab.clean .git/lab-hook.sh filter-clean
&lt;span class="nb"&gt;exit &lt;/span&gt;1

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The ownership check, and the runners that turn it off
&lt;/h2&gt;

&lt;p&gt;Since git 2.35.2 (the fix for &lt;a href="https://github.blog/open-source/git/git-security-vulnerability-announced/" rel="noopener noreferrer"&gt;CVE-2022-24765&lt;/a&gt;), git refuses to work in a repository owned by another user unless that path is allowed by &lt;code&gt;safe.directory&lt;/code&gt;. On the Pi, that check stopped the poisoned tar as soon as another user owned it.&lt;/p&gt;

&lt;p&gt;On GitHub-hosted runners, the same step ran everything. The script prints where &lt;code&gt;safe.directory&lt;/code&gt; comes from, and on the Ubuntu runner (image ubuntu24 20260927.320.1) it said:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git 2.55.0 on Linux x86_64
safe.directory already set on this machine: system file:/etc/gitconfig *
original repo copy ran: - then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0
git clone copy ran: nothing then ran: nothing exit 0/0
git clone from a bundle copy ran: nothing then ran: nothing exit 0/0
tar of the working copy copy ran: nothing then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0
same tar, owned by another user copy ran: - then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0
  ...with no system or global config copy ran: - then ran: nothing exit 128/129 fatal: detected dubious ownership in repository at '/tmp/tmp.Y1kIjtRPwM/other-owner'
  ...with -c safe.directory=* copy ran: - then ran: filter-clean,fsmonitor,hook:post-index-change exit 0/0

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Ubuntu image writes &lt;code&gt;directory = *&lt;/code&gt; into &lt;code&gt;/etc/gitconfig&lt;/code&gt;, and the macOS image (macos26 20260907.0351.1 in our run) adds it to the global config; you can see both in the &lt;a href="https://github.com/actions/runner-images/blob/main/images/ubuntu/scripts/build/install-git.sh" rel="noopener noreferrer"&gt;Ubuntu&lt;/a&gt; and &lt;a href="https://github.com/actions/runner-images/blob/main/images/macos/scripts/build/install-git.sh" rel="noopener noreferrer"&gt;macOS&lt;/a&gt; image scripts. The Ubuntu script's comment says why: git 2.35.2 "introduces security fix that breaks action\checkout". With system and global config switched off, git 2.55 refused the copy exactly like git 2.39 on the Pi, so the behaviour is git's and the wildcard is the image's.&lt;/p&gt;

&lt;p&gt;Two caveats keep this in proportion. The ownership check only helps when the files belong to a different user; a cache that the job restores as itself passes the check with or without a wildcard. And a refusal stops git, it does not make the repository safe. Still, on those runners the ownership check is not part of your defence, so it comes down to whether you restore &lt;code&gt;.git&lt;/code&gt; at all.&lt;/p&gt;

&lt;p&gt;To see what your own runners and build containers do, print every entry with its scope and file. Git only honours &lt;code&gt;safe.directory&lt;/code&gt; from system, global and command-line config, so a repo cannot allow itself. A &lt;code&gt;*&lt;/code&gt; in any of those means the check is off for all paths, unless a later empty entry resets the list, and removing a global &lt;code&gt;*&lt;/code&gt; does nothing if the system config still has one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git &lt;span class="nt"&gt;--no-pager&lt;/span&gt; config &lt;span class="nt"&gt;--show-scope&lt;/span&gt; &lt;span class="nt"&gt;--show-origin&lt;/span&gt; &lt;span class="nt"&gt;--get-all&lt;/span&gt; safe.directory

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What to do
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Do not restore .git from a cache or artifact
&lt;/h3&gt;

&lt;p&gt;Cache what is expensive to rebuild, not the repository. &lt;code&gt;.git&lt;/code&gt; rarely needs caching: actions/checkout fetches a single commit by default, and a fetch never carries config or hooks. If you need full history for speed, cache a bundle (&lt;code&gt;git bundle create&lt;/code&gt;) and clone from it; in our test, a bundle clone ran nothing.&lt;/p&gt;

&lt;p&gt;This closes one route, not cache poisoning in general. A restored &lt;code&gt;node_modules&lt;/code&gt; or provider directory can also contain code that runs, and GitHub's own &lt;a href="https://docs.github.com/en/actions/concepts/workflows-and-actions/dependency-caching" rel="noopener noreferrer"&gt;cache security guidance&lt;/a&gt; says to treat restored caches as untrusted input. Keep caches separated by trust level, so a pull request job cannot write what the main branch job restores.&lt;/p&gt;

&lt;p&gt;Treat artifacts the same way. If a later job downloads an artifact that contains a &lt;code&gt;.git&lt;/code&gt; directory and runs any git command inside it, that job trusts whoever produced the artifact.&lt;/p&gt;

&lt;h3&gt;
  
  
  Audit before the first git command
&lt;/h3&gt;

&lt;p&gt;When you cannot avoid a copied &lt;code&gt;.git&lt;/code&gt; (a synced folder, a devcontainer volume, a repo an agent has had write access to), read its config before you run git in it. &lt;code&gt;git --no-pager config --list&lt;/code&gt; reads config files without running any of the programs above; keep &lt;code&gt;--no-pager&lt;/code&gt;, because a configured pager can run when the output goes to a terminal.&lt;/p&gt;

&lt;p&gt;Our &lt;code&gt;scripts/audit.sh&lt;/code&gt; wraps that. It lists repo-level keys whose value is a command, a hooks directory or an include (including &lt;code&gt;hook.&amp;lt;name&amp;gt;.command&lt;/code&gt;, process filters, shell aliases and &lt;code&gt;pager.&amp;lt;cmd&amp;gt;&lt;/code&gt;), plus executable hooks in the repo's hooks directory, and exits 0 (none), 1 (found) or 2 (it could not inspect the repo, for example because git refused it). &lt;code&gt;scripts/audit-selftest.sh&lt;/code&gt; checks it against every fixture plus odd layouts, such as a driver name containing &lt;code&gt;=&lt;/code&gt;, a symlinked hook, a linked worktree and an include outside &lt;code&gt;.git&lt;/code&gt;, and confirms that auditing ran nothing. It is a detector for the settings it knows, not a safety certificate.&lt;/p&gt;

&lt;p&gt;Run a copy you trust, kept outside the directory you are checking; a poisoned workspace can replace any script inside it. In CI, that means fetching the audit at a pinned commit before you restore anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# before restoring the cache: fetch the audit from a pinned commit, outside the workspace&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$RUNNER_TEMP&lt;/span&gt;&lt;span class="s2"&gt;/audit.sh"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://raw.githubusercontent.com/The-DevOps-Daily/git-config-exec-check/&amp;lt;commit-sha&amp;gt;/scripts/audit.sh"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="c"&gt;# after restoring it&lt;/span&gt;
bash &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$RUNNER_TEMP&lt;/span&gt;&lt;span class="s2"&gt;/audit.sh"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$GITHUB_WORKSPACE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"restored repo can run programs"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expect some findings on developer machines. &lt;code&gt;git lfs install --local&lt;/code&gt; writes a filter driver into the repo config, and many teams use hooks on purpose. That is fine: the point is to see them before git runs them, in a place where you did not put them yourself.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you write tools that shell out to git
&lt;/h3&gt;

&lt;p&gt;GitLab's advice is to override every relevant key on every call, or avoid git's filter and textconv machinery. From our results, for the commands a tool runs to read state (&lt;code&gt;status&lt;/code&gt;, &lt;code&gt;diff&lt;/code&gt;, &lt;code&gt;blame&lt;/code&gt;, &lt;code&gt;log&lt;/code&gt;), that means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pass &lt;code&gt;-c core.fsmonitor=false -c core.hooksPath=/dev/null&lt;/code&gt; on every call, including &lt;code&gt;git status&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;On git 2.55, also pass &lt;code&gt;-c hook.&amp;lt;event&amp;gt;.enabled=false&lt;/code&gt; for each event your commands can fire. Our table saw &lt;code&gt;pre-commit&lt;/code&gt;, &lt;code&gt;post-commit&lt;/code&gt;, &lt;code&gt;post-checkout&lt;/code&gt;, &lt;code&gt;post-index-change&lt;/code&gt; and &lt;code&gt;reference-transaction&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;--no-ext-diff --no-textconv&lt;/code&gt; to diffs, and &lt;code&gt;--no-pager&lt;/code&gt; to anything that might reach a terminal.&lt;/li&gt;
&lt;li&gt;For filters, read the config first and clear &lt;code&gt;clean&lt;/code&gt;, &lt;code&gt;smudge&lt;/code&gt; and &lt;code&gt;process&lt;/code&gt; for every driver it defines, or use commit-to-commit diffs when they are enough.&lt;/li&gt;
&lt;li&gt;For &lt;code&gt;fetch&lt;/code&gt; and &lt;code&gt;push&lt;/code&gt;, the remote URL and &lt;code&gt;core.sshCommand&lt;/code&gt; both come from the copied config. Set the transport yourself (for example &lt;code&gt;-c core.sshCommand=ssh&lt;/code&gt; and an explicit URL) or do not let the tool talk to remotes from that copy.&lt;/li&gt;
&lt;li&gt;When the repo came from outside, run git as a user that does not own it, with no &lt;code&gt;safe.directory&lt;/code&gt; wildcard in system or global config.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These cover what we tested. They are not a complete list of every setting git can turn into a command, which is why the audit and the "do not copy &lt;code&gt;.git&lt;/code&gt;" advice come first.&lt;/p&gt;

&lt;p&gt;If you use one of the agents named in the two reports, update it. The DeepSeek-Reasonix fix is Studio 2.21.0 and npm 1.39.3. Manifold's post has a table of fixed versions; four of its eight findings were still unpatched when it was published.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we did not test
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Windows&lt;/strong&gt; , credential helpers and editors. Each of those needs a different trigger.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fetch that succeeds&lt;/strong&gt; , and a process filter that speaks the protocol. Both can run more than our rows show.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any specific agent.&lt;/strong&gt; These scripts test git. How a given agent calls git decides which rows of the table apply to it, and the two reports above are the place to look for that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every command and every state.&lt;/strong&gt; We picked 15 common commands and one kind of change. &lt;code&gt;git rev-parse&lt;/code&gt; and &lt;code&gt;git log --oneline&lt;/code&gt; ran nothing here; that is not a promise about other commands or other repo states, so run &lt;code&gt;matrix.sh&lt;/code&gt; with the commands your tools use.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have read our posts on &lt;a href="https://devops-daily.com/posts/pre-commit-hooks-security-guide" rel="noopener noreferrer"&gt;pre-commit hook security&lt;/a&gt; or the &lt;a href="https://devops-daily.com/posts/mcp-design-flaw-rce-supply-chain-risk" rel="noopener noreferrer"&gt;MCP design flaw&lt;/a&gt;, this is the same lesson from another side. The dangerous input is not always code you run on purpose. Sometimes it is the configuration of the tool you run, read from the directory you are standing in.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/poisoned-git-config-git-status-runs-code" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>git</category>
      <category>security</category>
      <category>cicd</category>
      <category>aiagents</category>
    </item>
    <item>
      <title>How Shopify Survives Black Friday: The Flash-Sale Playbook</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Tue, 06 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/how-shopify-survives-black-friday-the-flash-sale-playbook-4h17</link>
      <guid>https://dev.to/devopsdaily/how-shopify-survives-black-friday-the-flash-sale-playbook-4h17</guid>
      <description>&lt;p&gt;Shopify merchants sold &lt;a href="https://www.shopify.com/news/bfcm-data-2025" rel="noopener noreferrer"&gt;$14.6 billion over Black Friday Cyber Monday 2025&lt;/a&gt;, and the busiest minute, 12:01 p.m. EST on Black Friday, ran at $5.1 million in sales. A minute like that is not a surprise. It is on the calendar a year ahead, so the engineering problem is not reacting to a spike. It is rehearsing for one.&lt;/p&gt;

&lt;p&gt;Shopify's engineering blog describes the rehearsal in unusual detail: load tests at 150% of last year's peak, capacity planned with the cloud providers months ahead, a throttle at the edge that queues buyers when checkout is full, and inventory reserved during payment so two buyers cannot claim the last unit. This post walks through each from Shopify's own write-ups, then rebuilds the checkout half on a Raspberry Pi with k6 to measure what the last two steps buy.&lt;/p&gt;

&lt;p&gt;The short version: checking stock before payment oversold 41 to 47 units of 500 in three runs, and reserving the unit first sold exactly 500. When buyers arrived twice as fast as payment could take them, a plain queue charged more shoppers after they had given up than it confirmed purchases for. A simple waiting room kept checkout under 400 ms at p95 instead, at the cost of turning the overflow into shoppers who gave up waiting.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Shopify rehearses all year: its load generator, Genghis, runs checkout flows in production with flash-sale bursts on top, at 150% of last year's load, and capacity is planned months ahead.&lt;/li&gt;
&lt;li&gt;When checkout is full, an edge throttle queues buyers. The first version was a lottery that left some waiting 40 minutes; a signed first-attempt timestamp made it fair.&lt;/li&gt;
&lt;li&gt;Inventory is reserved when payment starts and claimed when it succeeds.&lt;/li&gt;
&lt;li&gt;In our runs the oversell from check-then-pay was roughly arrival rate times payment time: lower traffic shrank it but did not remove the race. At twice the payment capacity, the question was not how many orders got through but who got them.&lt;/li&gt;
&lt;li&gt;Our first batch of runs quietly skipped up to a fifth of its shoppers. Fail arrival-rate load tests on &lt;code&gt;dropped_iterations&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Comfort with HTTP, SQL transactions and the idea of a load test&lt;/li&gt;
&lt;li&gt;To run the demo: Node.js 22.13 or newer (for the built-in &lt;code&gt;node:sqlite&lt;/code&gt;) and the k6 binary; no Docker, no database server&lt;/li&gt;
&lt;li&gt;About 25 minutes of machine time for the full set of recorded runs&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A flash sale is an overload you can schedule
&lt;/h2&gt;

&lt;p&gt;Bart de Water, who worked on Shopify Payments, defined the term in his &lt;a href="https://www.infoq.com/presentations/shopify-architecture-flash-sale/" rel="noopener noreferrer"&gt;QCon talk on Shopify's flash-sale architecture&lt;/a&gt;: "A flash sale is a sale for a limited amount of time, often with limited stock. It's over in a flash because the product can sell out in seconds, even if there are thousands of items in inventory." He also draws the line that matters here: "Storefront is mostly about read traffic, while our checkout does most of the writing and has to interact with external systems as well."&lt;/p&gt;

&lt;p&gt;Reads cache well. Writes that call a payment provider do not, and that is where Shopify got hurt. &lt;a href="https://shopify.engineering/surviving-flashes-of-high-write-traffic-using-scriptable-load-balancers-part-i" rel="noopener noreferrer"&gt;The first of two posts on its checkout throttle&lt;/a&gt; describes Kylie Cosmetics running sales that sold out quickly, roughly every week, and one in February 2016 that "took down not just her store, but all others on the database shard where her store was allocated." What followed was a set of habits Shopify still writes about, and de Water's talk has the one-line reason they keep paying off: "today's flash sale will be tomorrow's base load."&lt;/p&gt;

&lt;p&gt;Those habits answer three questions. Will the platform hold at the peak? What happens to buyers who arrive when checkout is full? Does the one write that matters stay correct under contention?&lt;/p&gt;

&lt;h2&gt;
  
  
  Load testing as a discipline, not a launch task
&lt;/h2&gt;

&lt;p&gt;Shopify's &lt;a href="https://shopify.engineering/bfcm-readiness-2025" rel="noopener noreferrer"&gt;2025 BFCM readiness post&lt;/a&gt; opens with "Bimonthly fire drills all year, simulating 150% of last year's BFCM load." The tool is in-house:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Our load testing tool Genghis runs scripted workflows that mimic user behavior like browsing, cart adds, and checkout flows. We gradually ramp traffic to find breaking points. Tests run on production infrastructure simultaneously from three GCP regions (us-central, us-east, and europe-west4) to simulate global traffic patterns. We inject flash sale bursts on top of baseline load to test peak capacity.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Five major scale tests ran from April to October 2025. The fourth reached 146 million requests per minute and more than 80,000 checkouts per minute; the last went to the p99 forecast of 200 million requests per minute. Early tests found that "core operations threw errors and checkout queues backed up", and adding authenticated checkout "exposed rate-limit paths that anonymous browsing never touches."&lt;/p&gt;

&lt;p&gt;Three details worth copying at any scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Production-shaped targets.&lt;/strong&gt; De Water's talk describes Genghis hitting benchmark stores in production, at least one per pod, "at least weekly", paying through a benchmark gateway that "can respond with both successful and failed payments with a realistic distribution of response time latencies that we see in production." The demo copies that idea.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The burst, not just the volume.&lt;/strong&gt; A &lt;a href="https://shopify.engineering/scale-performance-testing" rel="noopener noreferrer"&gt;2023 post&lt;/a&gt; lists a flash-sale flow in which "simulated users purchase a single product", plus an "abort switch" that stops every test at once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Written-down failure.&lt;/strong&gt; &lt;a href="https://github.com/Shopify/toxiproxy" rel="noopener noreferrer"&gt;Toxiproxy&lt;/a&gt;, Shopify's open source tool for simulating network conditions, injects failures during load tests, and findings go into a Resiliency Matrix of failure scenarios, recovery objectives and runbooks.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Pick the load model before the tool
&lt;/h3&gt;

&lt;p&gt;What decides whether a load test can see a flash sale at all is the workload model. k6's page on &lt;a href="https://grafana.com/docs/k6/latest/using-k6/scenarios/concepts/open-vs-closed/" rel="noopener noreferrer"&gt;open and closed models&lt;/a&gt; states the trap: "When the target system is stressed and starts to respond more slowly, a closed model load test will wait, resulting in increased iteration durations and a tapering off of the arrival rate of new VU iterations." A closed test slows down exactly when the system does and reports a calm result. Buyers at a drop do not wait for the previous buyer's page to load, so you need an open model, where arrivals follow a schedule whatever the server is doing.&lt;/p&gt;

&lt;p&gt;All three common open source tools can drive an open model and fail a CI job; k6 and Gatling also have explicit closed-model options:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;k6&lt;/th&gt;
&lt;th&gt;Gatling&lt;/th&gt;
&lt;th&gt;Artillery&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Scripts&lt;/td&gt;
&lt;td&gt;JavaScript&lt;/td&gt;
&lt;td&gt;Java, JavaScript, Kotlin, Scala&lt;/td&gt;
&lt;td&gt;YAML, or JavaScript and TypeScript&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open model&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;constant-arrival-rate&lt;/code&gt;, &lt;code&gt;ramping-arrival-rate&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;constantUsersPerSec&lt;/code&gt;, &lt;code&gt;rampUsersPerSec&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;phases with &lt;code&gt;arrivalRate&lt;/code&gt; and &lt;code&gt;rampTo&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Closed model&lt;/td&gt;
&lt;td&gt;VU-based executors such as &lt;code&gt;constant-vus&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;constantConcurrentUsers&lt;/code&gt;, &lt;code&gt;rampConcurrentUsers&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;arrival-based phases; &lt;code&gt;maxVusers&lt;/code&gt; only caps concurrency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pass or fail&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://grafana.com/docs/k6/latest/using-k6/thresholds/" rel="noopener noreferrer"&gt;thresholds&lt;/a&gt;, non-zero exit&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.gatling.io/concepts/assertions/" rel="noopener noreferrer"&gt;assertions&lt;/a&gt;: "If at least one assertion fails, the simulation fails"&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.artillery.io/docs/reference/extensions/ensure" rel="noopener noreferrer"&gt;&lt;code&gt;ensure&lt;/code&gt; plugin&lt;/a&gt;, non-zero exit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On merit: k6 is a single Go binary whose thresholds work on custom metrics, which made "fail if any unit was sold twice" a one-line rule in the demo. Gatling separates &lt;a href="https://docs.gatling.io/concepts/injection/" rel="noopener noreferrer"&gt;open and closed injection&lt;/a&gt; in its API, so the model is explicit in every script, and suits JVM teams. Artillery's &lt;a href="https://www.artillery.io/docs/reference/test-script" rel="noopener noreferrer"&gt;YAML scenarios&lt;/a&gt; are the quickest to write for plain HTTP, and its Playwright engine drives real browsers when the checkout page's JavaScript matters. Any of them could run the demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pre-scale for the peak you know is coming
&lt;/h2&gt;

&lt;p&gt;Shopify's 2025 preparation started in March with capacity planning, including "submitting our estimates to our cloud providers so they don't run out of cloud." In the uncertain year of 2020, &lt;a href="https://shopify.engineering/capacity-planning-shopify" rel="noopener noreferrer"&gt;Capacity Planning at Scale&lt;/a&gt; records the choice: "We decided to scale to our more aggressive growth scenarios to ensure our platform is stable regardless of what happens." Before scale tests, components are brought up to a "BFCM profile" ahead of time.&lt;/p&gt;

&lt;p&gt;Change is managed the same way. A &lt;a href="https://shopify.engineering/preparing-shopify-for-black-friday-cyber-monday" rel="noopener noreferrer"&gt;2018 post&lt;/a&gt; describes a feature freeze that "starts several weeks before BFCM" and a code freeze a few days before. The 2025 post puts it as a rule: "We don't use BFCM as a release deadline." Architectural changes and migrations land months earlier.&lt;/p&gt;

&lt;p&gt;Why not let autoscaling handle it? Shopify's posts do not say they turn it off, and we found no primary source that does. But in the demo below the stock is gone about six seconds into the sale, while a Kubernetes Horizontal Pod Autoscaler checks metrics &lt;a href="https://kubernetes.io/docs/concepts/workloads/autoscaling/horizontal-pod-autoscale/" rel="noopener noreferrer"&gt;every 15 seconds by default&lt;/a&gt;, before a new pod is even scheduled. Autoscaling suits the slow tides of a long weekend; for the first minute of a drop, the capacity has to exist already. Our &lt;a href="https://devops-daily.com/exercises/kubernetes-hpa-lab" rel="noopener noreferrer"&gt;Kubernetes HPA lab&lt;/a&gt; and &lt;a href="https://devops-daily.com/games/scaling-simulator" rel="noopener noreferrer"&gt;horizontal vs vertical scaling simulator&lt;/a&gt; let you feel that lag safely.&lt;/p&gt;

&lt;p&gt;Some capacity cannot be pre-scaled at all because it belongs to someone else, like a payment provider. The second experiment is about what to do when arrivals exceed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put the queue at the edge, not inside checkout
&lt;/h2&gt;

&lt;p&gt;After the Kylie outage, Shopify had a week before the next sale. &lt;a href="https://shopify.engineering/surviving-flashes-of-high-write-traffic-using-scriptable-load-balancers-part-i" rel="noopener noreferrer"&gt;Part I&lt;/a&gt; explains why a plain rate limit would not do: "For customers anything that looked like the website crashing would be interpreted as such." The team built a throttle into its Nginx and OpenResty edge tier:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The solution we landed on was to throttle users using a leaky bucket algorithm built into our edge tier. [...] The number of requests served would be reset every period (in our case, 5 seconds), and it was up to us to inform the client when to retry a rejected request.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Buyers over the limit saw a queue page that polled &lt;code&gt;/checkout&lt;/code&gt;; those who got through received a signed cookie to skip the throttle for the session. The platform held, and then buyers complained of waiting up to 40 minutes for a 40-minute sale. The post admits that "in reality they were randomly polling the throttle", a lottery.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://shopify.engineering/surviving-flashes-of-high-write-traffic-using-scriptable-load-balancers-part-ii" rel="noopener noreferrer"&gt;Part II&lt;/a&gt; fixed fairness without adding state to the edge. A buyer's first checkout attempt is stamped with a timestamp in a signed cookie, and each load balancer computes a threshold, "the virtual version of the 'Now Serving: 42' counters at delis", that lets the earliest timestamps through to the leaky bucket. A feedback controller moves the threshold; after simulations, the team kept only its proportional term.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Buyer&lt;/strong&gt; clicks checkout&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge throttle&lt;/strong&gt; Nginx + Lua&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Checkout&lt;/strong&gt; writes + payment&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Outcomes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;under capacity: signed cookie, straight through&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;over capacity: queue page, poll, earliest timestamp first&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Five years later, de Water's &lt;a href="https://shopify.engineering/building-resilient-payment-systems" rel="noopener noreferrer"&gt;10 tips for resilient payment systems&lt;/a&gt; (2022) still describes "scriptable load balancers to throttle the amount of checkouts happening at any given time", with a waiting queue when demand exceeds capacity. It explains why with Little's Law and a sentence worth pinning above any capacity plan: "your application can't out scale the world." At the edge, a waiting buyer costs a cached page and a poll, not a worker, a database connection and a payment slot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Protect the one write that must not go wrong
&lt;/h2&gt;

&lt;p&gt;The 2025 post calls checkout, payment processing, order creation and fulfillment "critical journeys". Older parts of the architecture exist to keep them alive:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pods.&lt;/strong&gt; &lt;a href="https://shopify.engineering/a-pods-architecture-to-allow-shopify-to-scale" rel="noopener noreferrer"&gt;A pod&lt;/a&gt; "consists of a set of shops that live on a fully isolated set of datastores", which limits how far one shop's sale can reach. De Water adds that some extra-large merchants get a pod to themselves, and other systems stop a big merchant's flash sale from trying to "monopolize all the capacity" of a shared pod.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semian.&lt;/strong&gt; Shopify's &lt;a href="https://github.com/Shopify/semian" rel="noopener noreferrer"&gt;circuit breaker and bulkhead library for Ruby&lt;/a&gt; starts from the fact that slow resources fail slowly, and while threads wait on one, "the slow resource has caused a cascading failure by occupying workers and therefore losing capacity."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inventory reservations&lt;/strong&gt; , the part the demo rebuilds.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Shopify's May 2026 post, &lt;a href="https://shopify.engineering/scaling-inventory-reservations" rel="noopener noreferrer"&gt;We replaced Redis with MySQL for inventory reservations&lt;/a&gt;, says oversell protection works "by reserving inventory during payment processing", which it calls "a short hold that prevents two concurrent checkouts from claiming the same unit." Reserve holds items when payment starts; claim deducts them from the ledger when payment succeeds. The reservation used to be a Redis &lt;code&gt;DECR&lt;/code&gt;, but then the claim step had to update the MySQL ledger and clean up Redis, two operations that could not be wrapped in one atomic step. In MySQL, the obvious schema failed: "A single row with a quantity column couldn't handle the contention." The shipped design uses one row per sellable unit, taken with &lt;code&gt;SELECT ... FOR UPDATE SKIP LOCKED&lt;/code&gt; so concurrent checkouts take different rows, from a pool capped at 1,000 rows per item and location that a replenishment process refills. The post also ties this article's two halves together: "Slow reservations trigger throttling and a worse buyer experience." Remember the single-row detail; the demo's fix is the single-row version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rebuilding the checkout half on a Raspberry Pi
&lt;/h2&gt;

&lt;p&gt;The demo is a small checkout service and one k6 scenario, built to answer two questions a skeptical reader can check: does reserving before payment matter at modest traffic, and what does a queue in front of checkout change when arrivals exceed what payment can process?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/The-DevOps-Daily/flash-sale-playbook" rel="noopener noreferrer"&gt;The-DevOps-Daily/flash-sale-playbook on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Everything ran on a 4-core Raspberry Pi 4 that was also doing other work, with k6 2.3.0 and the service on the same machine, so absolute numbers are small. Every variant ran the same scenario on the same machine, interleaved with the others, and each run records the load average before and after it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The service&lt;/strong&gt; is one Node.js 24 process, one product, and an SQLite database in memory-backed storage. One process stands in for a fleet: the &lt;code&gt;await&lt;/code&gt; between the read and the write lets other requests run in between, as requests on separate servers would.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The payment provider&lt;/strong&gt; is a stand-in that takes 100 to 400 ms (uniform, seeded per run) and declines 5%. In the second experiment it also has a fixed number of slots.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The scenario&lt;/strong&gt; is a k6 &lt;code&gt;ramping-arrival-rate&lt;/code&gt; executor with every VU created before the sale. Each iteration is one shopper. After the sale, &lt;code&gt;teardown()&lt;/code&gt; asks the server what it sold and reports it as metrics, so thresholds fail the run on correctness, not just latency.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// k6/flash-sale.js, trimmed&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;scenarios&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;flash_sale&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ramping-arrival-rate&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;startRate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;timeUnit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;1s&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;preAllocatedVUs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vus&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;maxVUs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vus&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;stages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// drop: 0 -&amp;gt; 300/s over 10s, hold 20s, down over 5s&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;thresholds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;dropped_iterations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;count==0&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="c1"&gt;// k6 kept its own schedule&lt;/span&gt;
    &lt;span class="na"&gt;oversold_units&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;count==0&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="c1"&gt;// reported by teardown() from the server&lt;/span&gt;
    &lt;span class="na"&gt;orders_to_clients_that_left&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;count==0&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http_req_duration{name:checkout}&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;p(95)&amp;lt;2000&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before the recorded runs, &lt;code&gt;scripts/predict.mjs&lt;/code&gt; simulated every variant 400 times with an idealized copy of k6's arrival schedule (it has one extra arrival at the very end) and the same payment model, ignoring server CPU time. Its output from the start of the first batch, kept with the discarded runs, is identical to the one used here. Every recorded order count below fell inside the simulated 5th to 95th percentile range; a few latency figures landed just outside it, which is unsurprising for a model with no CPU time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Experiment 1: check, pay, then write
&lt;/h3&gt;

&lt;p&gt;The naive checkout is the order most people write first (simplified from &lt;code&gt;server.mjs&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;checkoutNaive&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shopper&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;stock&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;readStock&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;SKU&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// SELECT stock ...&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;409&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt; &lt;span class="c1"&gt;// sold out&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;authorizePayment&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// 100-400 ms; other requests run here&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payment&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;declined&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;402&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;decrement&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;SKU&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// UPDATE inventory SET stock = stock - 1&lt;/span&gt;
  &lt;span class="nf"&gt;recordOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shopper&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix moves a conditional decrement in front of payment, so the database decides who gets each unit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;checkoutReserve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shopper&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// UPDATE inventory SET stock = stock - 1 WHERE sku = ? AND stock &amp;gt; 0&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;changes&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reserve&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;SKU&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;changes&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;409&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt; &lt;span class="c1"&gt;// sold out, no payment call&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;authorizePayment&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payment&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;paid&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;release&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;SKU&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// give the unit back&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;402&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nf"&gt;recordOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shopper&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;drop&lt;/code&gt; profile ramps from zero to 300 new shoppers a second over 10 seconds, holds for 20 and ramps down over 5: 8,249 shoppers for 500 units. Each mode ran three times, interleaved to spread background load across both:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Checkout&lt;/th&gt;
&lt;th&gt;Orders&lt;/th&gt;
&lt;th&gt;Oversold&lt;/th&gt;
&lt;th&gt;Stock counter after&lt;/th&gt;
&lt;th&gt;Checkout p95&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;drop, seed 1&lt;/td&gt;
&lt;td&gt;naive&lt;/td&gt;
&lt;td&gt;544&lt;/td&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;td&gt;-44&lt;/td&gt;
&lt;td&gt;190 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drop, seed 2&lt;/td&gt;
&lt;td&gt;naive&lt;/td&gt;
&lt;td&gt;541&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;-41&lt;/td&gt;
&lt;td&gt;174 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drop, seed 3&lt;/td&gt;
&lt;td&gt;naive&lt;/td&gt;
&lt;td&gt;547&lt;/td&gt;
&lt;td&gt;47&lt;/td&gt;
&lt;td&gt;-47&lt;/td&gt;
&lt;td&gt;186 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drop, seed 1&lt;/td&gt;
&lt;td&gt;reserve&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;257 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drop, seed 2&lt;/td&gt;
&lt;td&gt;reserve&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;165 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drop, seed 3&lt;/td&gt;
&lt;td&gt;reserve&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;168 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;low, seed 1&lt;/td&gt;
&lt;td&gt;naive&lt;/td&gt;
&lt;td&gt;508&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;-8&lt;/td&gt;
&lt;td&gt;369 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;low, seed 1&lt;/td&gt;
&lt;td&gt;reserve&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;372 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every run started all its scheduled shoppers and used the same code, stamped by hash in its &lt;code&gt;run.json&lt;/code&gt;. The p95 covers every checkout request, including the quick &lt;code&gt;409&lt;/code&gt; sold-out answers most shoppers got.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Units sold beyond the 500 in stock&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;predicted mean&lt;/th&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Drop, naive, seed 1&lt;/td&gt;
&lt;td&gt;44 units&lt;/td&gt;
&lt;td&gt;42.5 units&lt;/td&gt;
&lt;td&gt;Naive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drop, naive, seed 2&lt;/td&gt;
&lt;td&gt;41 units&lt;/td&gt;
&lt;td&gt;42.5 units&lt;/td&gt;
&lt;td&gt;Naive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drop, naive, seed 3&lt;/td&gt;
&lt;td&gt;47 units&lt;/td&gt;
&lt;td&gt;42.5 units&lt;/td&gt;
&lt;td&gt;Naive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drop, reserve, seeds 1-3&lt;/td&gt;
&lt;td&gt;0 units&lt;/td&gt;
&lt;td&gt;0 units&lt;/td&gt;
&lt;td&gt;Reserve&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Low, naive, seed 1&lt;/td&gt;
&lt;td&gt;8 units&lt;/td&gt;
&lt;td&gt;6.6 units&lt;/td&gt;
&lt;td&gt;Naive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Low, reserve, seed 1&lt;/td&gt;
&lt;td&gt;0 units&lt;/td&gt;
&lt;td&gt;0 units&lt;/td&gt;
&lt;td&gt;Reserve&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Recorded k6 runs on a Raspberry Pi 4. Tick marks show the mean of 400 simulated runs from scripts/predict.mjs. Reserve sold exactly 500 every time.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;flash-sale-playbook&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# two of the recorded runs, as scripts/run-all.sh started them
$ scripts/run.sh drop-naive-1 drop MODE=naive SEED=1
drop-naive-1: k6 exit 99, results in results/drop-naive-1
$ scripts/run.sh drop-reserve-1 drop MODE=reserve SEED=1
drop-reserve-1: k6 exit 0, results in results/drop-reserve-1
# what the server says it sold
$ jq -c '{mode: .config.mode, stockInitial, orders, oversold, stockNow}' results/drop-naive-1/server-stats.json
{"mode":"naive","stockInitial":500,"orders":544,"oversold":44,"stockNow":-44}
$ jq -c '{mode: .config.mode, stockInitial, orders, oversold, stockNow}' results/drop-reserve-1/server-stats.json
{"mode":"reserve","stockInitial":500,"orders":500,"oversold":0,"stockNow":0}
# why the naive run exited 99
$ grep -A14 THRESHOLDS results/drop-naive-1/k6-output.txt
  █ THRESHOLDS 

    dropped_iterations
    ✓ 'count==0' count=0

    http_req_duration{name:checkout}
    ✓ 'p(95)&amp;lt;2000' p(95)=189.78ms

    orders_to_clients_that_left
    ✓ 'count==0' count=0

    oversold_units
    ✗ 'count==0' count=44

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every naive drop run sold units that did not exist, 8 to 9% more than the stock, and left the counter negative, so nothing downstream would have stopped those orders. Every reservation run sold exactly 500. Latency could not tell them apart (most shoppers arrived after the sellout and got a quick &lt;code&gt;409&lt;/code&gt;), so a load test that only checked p95 and errors would have passed both.&lt;/p&gt;

&lt;p&gt;The size of the oversell is predictable. Every shopper who passes the stock check before the counter reaches zero gets an order, and the counter only moves when a payment completes. So at the moment it hits zero, about (arrival rate × average payment time × share of payments approved) shoppers are still in flight. The 500th naive order landed about 6.2 seconds after the sale opened, with arrivals near 185 a second: 185 × 0.25 s × 0.95 ≈ 44. The simulation predicted a mean of 42.5, with 5th to 95th percentiles of 37 and 48; the runs landed at 44, 41 and 47.&lt;/p&gt;

&lt;p&gt;Lower traffic shrinks it but does not remove the race: at a tenth of the arrival rate, the prediction was 4 to 9 and the run oversold 8 (1.6% of the stock). A slower payment step makes it larger. The bug is about how many payments are in flight when the last unit goes, not about Black Friday traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Experiment 2: more buyers than payment can take
&lt;/h3&gt;

&lt;p&gt;The second scenario is the BFCM shape: plenty of stock, arrivals beyond what checkout can process. Payment has 8 slots and averages 250 ms, so it completes about 32 checkouts a second (Little's Law: 8 / 0.25 s). The &lt;code&gt;peak&lt;/code&gt; profile ramps to 64 shoppers a second, twice that, holds for 30 seconds and ramps down over 10: 2,559 shoppers. Each will spend at most 15 seconds trying to buy, in one slow request or several tries.&lt;/p&gt;

&lt;p&gt;Three variants, one change each, on the reservation code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Queue&lt;/strong&gt; : no protection. Requests wait for a payment slot as long as it takes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deadline&lt;/strong&gt; : the baseline a hostile reader would ask for. k6 sends the time the shopper gives up, and the server skips the charge (&lt;code&gt;504&lt;/code&gt;) when less than the stand-in's maximum payment time, 400 ms, is left.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Waiting room&lt;/strong&gt; : at most 8 checkouts in progress; everyone else gets &lt;code&gt;503&lt;/code&gt; with &lt;code&gt;Retry-After: 1&lt;/code&gt; and retries after a jittered half to one and a half seconds, giving up when the next wait would take them past 15 seconds. This is Shopify's random-polling first version, the lottery, not the timestamp fix.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Three runs per variant, seeds 1 to 3&lt;/th&gt;
&lt;th&gt;Queue&lt;/th&gt;
&lt;th&gt;Deadline&lt;/th&gt;
&lt;th&gt;Waiting room&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shopper saw a confirmation&lt;/td&gt;
&lt;td&gt;1,028 to 1,071&lt;/td&gt;
&lt;td&gt;1,794 to 1,834&lt;/td&gt;
&lt;td&gt;1,603 to 1,642&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shopper saw a timeout&lt;/td&gt;
&lt;td&gt;1,437 to 1,481&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;14 to 18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shopper was refused at the deadline&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;636 to 670&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shopper gave up after "please wait"&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;826 to 847&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shopper saw a decline&lt;/td&gt;
&lt;td&gt;50 to 61&lt;/td&gt;
&lt;td&gt;89 to 96&lt;/td&gt;
&lt;td&gt;77 to 91&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server charged a shopper who had left&lt;/td&gt;
&lt;td&gt;1,375 to 1,408&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;10 to 14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Successful checkout request, p95&lt;/td&gt;
&lt;td&gt;14.2 to 14.3 s&lt;/td&gt;
&lt;td&gt;14.9 s&lt;/td&gt;
&lt;td&gt;387 to 395 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to purchase, median&lt;/td&gt;
&lt;td&gt;6.4 to 6.7 s&lt;/td&gt;
&lt;td&gt;12.5 to 13.1 s&lt;/td&gt;
&lt;td&gt;4.3 to 5.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Most checkouts in progress at once&lt;/td&gt;
&lt;td&gt;1,122 to 1,142&lt;/td&gt;
&lt;td&gt;948 to 949&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Waiting-room responses served&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;21,892 to 22,188&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Predicted confirmations, mean&lt;/td&gt;
&lt;td&gt;1,048&lt;/td&gt;
&lt;td&gt;1,809&lt;/td&gt;
&lt;td&gt;1,626&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Predicted charges after leaving, mean&lt;/td&gt;
&lt;td&gt;1,385&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first five rows add up to all 2,559 shoppers in every run. The charges in the sixth row overlap with the timeouts: in the queue variant, most shoppers who saw a timeout were charged anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2,559 shoppers, payment capacity about 32 checkouts a second&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;predicted mean&lt;/th&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Queue: confirmed purchases&lt;/td&gt;
&lt;td&gt;1049 shoppers&lt;/td&gt;
&lt;td&gt;1047.9 shoppers&lt;/td&gt;
&lt;td&gt;Confirmed purchase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Queue: charged after giving up&lt;/td&gt;
&lt;td&gt;1389 shoppers&lt;/td&gt;
&lt;td&gt;1384.8 shoppers&lt;/td&gt;
&lt;td&gt;Charged after giving up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deadline: confirmed purchases&lt;/td&gt;
&lt;td&gt;1810 shoppers&lt;/td&gt;
&lt;td&gt;1809.5 shoppers&lt;/td&gt;
&lt;td&gt;Confirmed purchase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deadline: charged after giving up&lt;/td&gt;
&lt;td&gt;0 shoppers&lt;/td&gt;
&lt;td&gt;0 shoppers&lt;/td&gt;
&lt;td&gt;Charged after giving up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Waiting room: confirmed purchases&lt;/td&gt;
&lt;td&gt;1622 shoppers&lt;/td&gt;
&lt;td&gt;1626.1 shoppers&lt;/td&gt;
&lt;td&gt;Confirmed purchase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Waiting room: charged after giving up&lt;/td&gt;
&lt;td&gt;12 shoppers&lt;/td&gt;
&lt;td&gt;16.2 shoppers&lt;/td&gt;
&lt;td&gt;Charged after giving up&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Mean of three recorded k6 runs per variant on a Raspberry Pi 4. Tick marks show the mean of 400 simulated runs from scripts/predict.mjs.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Through the 45 seconds of overload, every variant completed 29 to 31 orders a second, by the server's own timestamps. Payment was the limit and none of the three changed it. What changed was who got those orders, and how long everyone waited.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The queue&lt;/strong&gt; turned the overload into charges nobody saw. The line for payment grew past 1,100 requests; once the wait passed 15 seconds, nearly every request was served after its shopper had given up. Of about 2,440 orders per run, roughly 1,390 went to shoppers whose browser had shown a timeout, more than the 1,050 or so who saw a confirmation. In a real shop, each is a card charged for a purchase the buyer thinks failed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The deadline&lt;/strong&gt; removed those charges and sold the most, about 1,810. The cost is time: the queue still settles near the patience limit, so the median buyer waited 12.5 to 13.1 seconds, with up to 949 requests open at once. In a thread-per-request server each would also hold a worker, the failure Semian exists to prevent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The waiting room&lt;/strong&gt; kept checkout fast: 8 in progress at most, under 400 ms at p95, and a median time to purchase of 4.3 to 5.0 seconds including the waiting. Between 826 and 847 shoppers per run gave up after close to 15 seconds of "please wait", and the server answered about 22,000 cheap &lt;code&gt;503&lt;/code&gt;s, around 8.6 per shopper. That is the trade Shopify made by moving the wait to a cached page at the edge.&lt;/p&gt;

&lt;p&gt;It confirmed about 10% fewer purchases than the deadline queue, and the timestamps show where most of that gap built up. Both sold about 30 a second through the overload. After arrivals stopped, 50 seconds into the sale, the deadline queue kept serving its newest arrivals, who still had patience left, until about 62 seconds; the waiting room ran dry at about 58. Server-recorded orders after the 50-second mark account for about 154 of the roughly 175-order gap. The cause is this particular polling policy rather than waiting rooms in general: admission is random, so shoppers who had waited longest gave up first; a shopper gives up as soon as the next random wait would cross the 15-second limit, up to 1.5 seconds early; and a freed slot sits idle until someone's retry arrives. It is Shopify's Part I lesson in miniature: random retries spend the patience of the people who came first. The 10 to 14 late charges are shoppers admitted with less patience left than their payment took. Adding the deadline check to the waiting room should remove most of them; we did not run that combination.&lt;/p&gt;

&lt;h3&gt;
  
  
  The bug in our own load test
&lt;/h3&gt;

&lt;p&gt;Our first full batch of runs, kept under &lt;code&gt;results/discarded/&lt;/code&gt;, started with 50 pre-allocated VUs and let k6 create more on demand. In the peak runs every waiting shopper holds a VU; k6 could not create them fast enough and skipped the shoppers it could not start on time: 541, 542 and 440 of 2,559. The runs completed, the summaries looked plausible, and the queue variant looked far better than it should, because it had faced a lighter sale.&lt;/p&gt;

&lt;p&gt;k6 reports this as &lt;code&gt;dropped_iterations&lt;/code&gt;, and our summary script flagged it. The fix was to create every VU up front and fail the run on &lt;code&gt;dropped_iterations: ['count==0']&lt;/code&gt;. Later, a burst of unrelated work on the Pi pushed the load average near 9 on four cores. One peak run in that window skipped 253 shoppers, and another started 2,557 of 2,559 without reporting any as dropped, which only surfaced when the summary script compared runs. Both were re-run once the machine was quiet and kept with a note. The rule for throwing a run away was incomplete delivery of the schedule, not load: &lt;code&gt;peak-queue-3&lt;/code&gt; started under a load average of 8.9, delivered all 2,559 shoppers, landed inside its predicted ranges and was kept. Before you believe anything an arrival-rate test says, check that it delivered its schedule.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not show
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scale.&lt;/strong&gt; A few hundred requests a second on one small machine, with k6 and the server sharing four cores. The direction of each result should hold; the absolute numbers do not transfer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hot-row contention.&lt;/strong&gt; SQLite serializes writes, so the single-row conditional &lt;code&gt;UPDATE&lt;/code&gt; never fought for a row lock the way it would in MySQL or Postgres. Shopify says that design "couldn't handle the contention" at its scale. The demo shows that reserving before payment is necessary, not that one row is enough.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The cost of a held request.&lt;/strong&gt; Node.js holds a waiting connection cheaply. In a thread-per-request server, the queue and deadline variants would look worse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fair waiting room.&lt;/strong&gt; Ours is the lottery Shopify replaced. A timestamp-ordered queue should change who gets served and narrow the spread of waits; we did not build one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real payment latency.&lt;/strong&gt; The stand-in and the predictor use a uniform 100 to 400 ms with no long tail. Real providers have tails, which make every queue worse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rare cases.&lt;/strong&gt; Three runs per variant (one for &lt;code&gt;low&lt;/code&gt;) agree with each other, and their order counts sit inside the simulated ranges, but three runs cannot rule out an occasional bad one.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The playbook, condensed
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Rehearse the shape, not just the volume.&lt;/strong&gt; Model the drop as a burst on baseline traffic with an open-model load test, and run it on a schedule.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make the load test check correctness.&lt;/strong&gt; After the run, ask the system what it sold. Fail on oversells, on charges to buyers who had left and on dropped arrivals, not only on p95.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Have the capacity before the minute you need it.&lt;/strong&gt; Forecast, agree capacity with providers, scale up ahead of the peak and stop risky changes well before it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bound the work that reaches checkout.&lt;/strong&gt; Admit what payment and inventory can finish and park everyone else somewhere cheap, ideally at the edge and in arrival order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reserve before you charge, and release on failure.&lt;/strong&gt; The database decides who gets the last unit, not the application's memory of an earlier read.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Shopify's own posts are worth reading in full, starting with the &lt;a href="https://shopify.engineering/bfcm-readiness-2025" rel="noopener noreferrer"&gt;2025 readiness program&lt;/a&gt;, the &lt;a href="https://shopify.engineering/surviving-flashes-of-high-write-traffic-using-scriptable-load-balancers-part-i" rel="noopener noreferrer"&gt;checkout throttle story&lt;/a&gt; and the &lt;a href="https://shopify.engineering/scaling-inventory-reservations" rel="noopener noreferrer"&gt;inventory reservations post&lt;/a&gt;. For the neighbouring problems, our post on &lt;a href="https://dev.to/devopsdaily/how-stripe-avoids-double-charging-anyone-2me1"&gt;how Stripe avoids double-charging&lt;/a&gt; covers retries around the payment call, and the &lt;a href="https://devops-daily.com/games/rate-limit-simulator" rel="noopener noreferrer"&gt;rate limit simulator&lt;/a&gt; shows how leaky and token buckets behave under bursts.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/how-shopify-survives-black-friday-flash-sale-playbook" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>loadtesting</category>
      <category>k6</category>
      <category>scalability</category>
    </item>
    <item>
      <title>Touch Grass: 6 Project Ideas for People Who Live in a Terminal (Hacktoberfest DEV Challenge, Week 1)</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Mon, 05 Oct 2026 20:01:41 +0000</pubDate>
      <link>https://dev.to/devopsdaily/touch-grass-6-project-ideas-for-people-who-live-in-a-terminal-hacktoberfest-dev-challenge-week-1-1e7n</link>
      <guid>https://dev.to/devopsdaily/touch-grass-6-project-ideas-for-people-who-live-in-a-terminal-hacktoberfest-dev-challenge-week-1-1e7n</guid>
      <description>&lt;p&gt;The theme for Week 1 of the Hacktoberfest DEV Challenge is out, and it is a little personal for anyone in DevOps: &lt;strong&gt;Touch Grass&lt;/strong&gt;. Build something with open-weight models or open-source AI that gets people off the screen and into the world. Submissions are due by &lt;strong&gt;October 11, 2026 at 11:59 PM PDT&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The organisers put it well: the best builds should make the screen the shortest part of the experience. That is a fun constraint for people whose job is mostly screens. This post gives you six project ideas with a DevOps flavour. Each one is small enough for a week of evenings, gets someone outside, and has a clear answer to the question the judges ask in every entry: &lt;strong&gt;why does open matter for what you built?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules in 30 seconds
&lt;/h2&gt;

&lt;p&gt;From the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;challenge page&lt;/a&gt; and the &lt;a href="https://dev.to/devteam/join-the-hacktoberfest-open-source-ai-challenge-week-1-touch-grass-2450-in-prizes-across-17-4pom"&gt;Week 1 launch post&lt;/a&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The prompt:&lt;/strong&gt; build something with open-source AI at its core: an open-weight model, an open-source agent harness or framework, local inference, or all three. The open pieces should be what makes the project work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The theme:&lt;/strong&gt; get people off the screen and into the world. Hiking, gardening, birding, run clubs, fall foliage: if it gets someone outside, it counts. Bonus points if you take it outside, use it, and tell us how it went.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A new project,&lt;/strong&gt; built during the challenge. Pull requests to existing projects do not count this year.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The write-up matters most.&lt;/strong&gt; Writing quality has the biggest weight in judging. Explain why open innovation matters for what you built: does it run with no internet, keep data off someone else's server, let you swap or fine-tune the model, cost nothing to run?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One submission per challenge.&lt;/strong&gt; It is in the running for the overall prize and every partner category it qualifies for, but it can win only once per challenge. To enter a partner category, use that partner's technology and list the category in the Prize Categories section of your post. Missed the weekend challenge? Every week is a fresh start.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  1. The pager-safe walk, for your on-call friend
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; on-call engineers stay inside because they are afraid of missing a page. A whole week indoors, for alerts that mostly need nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does:&lt;/strong&gt; answers one question before you leave: "How far can I walk and still be back at my laptop within 15 minutes?" It draws that area on a map, so the walk is planned in seconds and the rest is outside. While you walk, a small watcher reads incoming alerts and sends one short verdict to your phone: "stay out" or "head back, this one looks real". Plain rules decide whether a page needs you; the model only explains why in one sentence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why open matters:&lt;/strong&gt; alert payloads are full of hostnames, customer names and internal URLs. With a local model, they never leave your machine. The walking area comes from open map data: OpenStreetMap plus an open-source routing engine such as &lt;a href="https://valhalla.github.io/valhalla/api/isochrone/api-reference/" rel="noopener noreferrer"&gt;Valhalla&lt;/a&gt;, which has an isochrone API for exactly this "how far in N minutes" question. Use pedestrian costing, and show the area as an estimate: hills and crossings make the real walk back slower.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Week scope:&lt;/strong&gt; one isochrone request for your home address, one map page, one watcher that reads a webhook or a JSON file of alerts. Skip live paging integrations; a sample alert file is enough for the demo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Partner fits:&lt;/strong&gt; Temporal (a watcher that must not lose an alert is a good fit for a durable workflow), Mastra (orchestrate the steps over an open model).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background:&lt;/strong&gt; our &lt;a href="https://devops-daily.com/posts/on-call-rotation-escalation-policy-guide" rel="noopener noreferrer"&gt;on-call rotation and escalation guide&lt;/a&gt; explains which pages should wake a human, which is exactly the rule set your watcher needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Field notes from your camera roll, for the friend who hikes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; after a hike, the photos sit unsorted on a phone, and the "what was that plant?" question never gets answered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does:&lt;/strong&gt; you drop a folder of hike photos on your laptop. A local vision model writes field notes: what each photo shows, with a confidence level, and the place it was taken (from the photo's GPS data). The result is a one-page trail log with a map. The screen time is five minutes after the hike, not during it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why open matters:&lt;/strong&gt; photos with GPS coordinates show where you live and where you walk. With an open-weight vision model running locally (Gemma 4 takes text and image input), the photos never leave your laptop. If you add hosted search, say in the post that only the notes and embeddings go up, and drop the exact coordinates first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Week scope:&lt;/strong&gt; one folder in, one Markdown or HTML page out. Read EXIF with an existing library. Ask the model to say "not sure" instead of guessing a species, and show its uncertainty in the page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Partner fits:&lt;/strong&gt; Gemma (run it locally), Tiger Data or MongoDB Atlas (store notes and embeddings so you can search "all the mushrooms from September").&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background:&lt;/strong&gt; none needed, just a recent hike. If you have not been on one, that is the first step of the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. A cron job for going outside, for everyone on your team
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; "I'll go for a walk later" never happens, because later is a meeting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does:&lt;/strong&gt; reads your calendar (an &lt;code&gt;.ics&lt;/code&gt; export), the weather forecast, and today's sunset time, and books the best 30 to 45 minutes outside before it gets dark. Then it learns from your history which suggestions you actually took. If you never go out at 13:00 on Mondays, it stops suggesting that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why open matters:&lt;/strong&gt; your calendar shows who you meet and when. It stays on your machine. The weather comes from &lt;a href="https://open-meteo.com/en/docs" rel="noopener noreferrer"&gt;Open-Meteo&lt;/a&gt;, which is free for non-commercial use with no API key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Week scope:&lt;/strong&gt; a script that prints today's best remaining window and writes it back to your calendar as an event. The learning part can be a CSV of past suggestions with "taken" or "skipped".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Partner fits:&lt;/strong&gt; TabPFN (its category is about predicting from a historical CSV, which is exactly your "taken or skipped" log), Gemma (write the suggestion in a friendly sentence).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background:&lt;/strong&gt; if cron syntax still surprises you, try the &lt;a href="https://devops-daily.com/games/cron-expression-simulator" rel="noopener noreferrer"&gt;cron expression simulator&lt;/a&gt; before you schedule anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Alerting for tomatoes, for a friend with a garden
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; garden apps either nag every day or stay quiet until the plants are dead. It is the same alert fatigue we fight at work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does:&lt;/strong&gt; an Arduino UNO Q in the garden reads soil moisture and temperature, combines it with the local frost forecast, and pages you only when the garden needs you: "frost tonight, cover the peppers" or "the bed by the fence has been dry for three days". Everything else goes into a weekly summary. Apply the rule we use for production: page on symptoms, not on every reading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why open matters:&lt;/strong&gt; the model runs on the board itself, with no cloud account and no subscription. The network is only needed to fetch the daily frost forecast and to send the message; the readings and the decisions stay on the board.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Week scope:&lt;/strong&gt; one sensor, one plant bed, two alert rules, one weekly summary. A soil moisture sensor and a frost check are enough for a strong demo video.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Partner fits:&lt;/strong&gt; Arduino (run a model on the UNO Q, sense and act), DigitalOcean (host the weekly summary page, if you want one outside your home network).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background:&lt;/strong&gt; our guide to &lt;a href="https://devops-daily.com/posts/slos-slis-error-budgets-practical-guide" rel="noopener noreferrer"&gt;SLOs, SLIs and error budgets&lt;/a&gt; is about services, but the same thinking decides when a tomato deserves a page.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. A trail buddy that works with no signal, for your homelab friend
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; the best trails have no mobile signal, and that is exactly when you want to know how far it is to water, or when you must turn back to beat the sunset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does:&lt;/strong&gt; before the hike, it downloads what it needs: the route, water sources and shelters from OpenStreetMap (through the &lt;a href="https://wiki.openstreetmap.org/wiki/Overpass_API" rel="noopener noreferrer"&gt;Overpass API&lt;/a&gt;), and the sunset time. On the trail, a Raspberry Pi or an old phone answers questions by voice: "How far to the next water?" "When do I need to turn back?" The distances and times come from plain code; the model only turns them into a short spoken answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why open matters:&lt;/strong&gt; it has to work with no internet at all. A small local model, local speech recognition such as &lt;a href="https://github.com/ggml-org/whisper.cpp" rel="noopener noreferrer"&gt;whisper.cpp&lt;/a&gt;, and local speech output such as &lt;a href="https://github.com/OHF-Voice/piper1-gpl" rel="noopener noreferrer"&gt;Piper&lt;/a&gt; make that possible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Week scope:&lt;/strong&gt; one trail, three questions, voice in and voice out. Test it in airplane mode before you leave home, then test it on the trail.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Partner fits:&lt;/strong&gt; Gemma (the local model), SerpApi (a pre-trip check for trail closures and news while you still have signal).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background:&lt;/strong&gt; if you build this on a Pi, the &lt;a href="https://devops-daily.com/games/linux-terminal" rel="noopener noreferrer"&gt;Linux terminal simulator&lt;/a&gt; is a quick refresher on the commands you will need over SSH.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The walking standup, for a remote team
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; remote teams spend the whole day in calls. The daily standup could be the one meeting you take outside.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does:&lt;/strong&gt; everyone records a 60-second voice note on their walk: what they did, what they will do, what blocks them. Back at the desk, a local model transcribes the notes and writes the standup summary for the team channel. Nobody has to look at a screen during the meeting itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why open matters:&lt;/strong&gt; team updates mention customers, incidents and people. With local transcription and a local model on a team server you control, they stay inside the team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Week scope:&lt;/strong&gt; a folder of audio files in, one summary out. Use your own notes for a week, then try it with a friend's team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Partner fits:&lt;/strong&gt; DigitalOcean (run the team server or an open-weight model on a GPU Droplet), ElevenLabs (transcription, if you are fine with sending audio to a hosted service, or a narrated version of the summary for the demo), Mastra (orchestrate transcription, summary and posting).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background:&lt;/strong&gt; keep the summary short and factual. The rule from our on-call guide applies here too: if nobody acts on a message, it should not be sent.&lt;/p&gt;

&lt;h2&gt;
  
  
  A starter for the weather-aware ideas
&lt;/h2&gt;

&lt;p&gt;Ideas 3, 4 and 5 all start with the same two pieces: some facts from an open data source, and a local model that turns them into one useful sentence. With &lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt;, Gemma 4 and Open-Meteo:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull gemma4:e4b

&lt;span class="c"&gt;# today's sunset and today's hourly rain chance (example coordinates: Berlin)&lt;/span&gt;
&lt;span class="nv"&gt;WEATHER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; &lt;span class="s2"&gt;"https://api.open-meteo.com/v1/forecast?latitude=52.52&amp;amp;longitude=13.41&amp;amp;hourly=precipitation_probability&amp;amp;daily=sunset&amp;amp;timezone=auto&amp;amp;forecast_days=1"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="nv"&gt;NOW&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nv"&gt;TZ&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Europe/Berlin &lt;span class="nb"&gt;date&lt;/span&gt; +%H:%M&lt;span class="si"&gt;)&lt;/span&gt;

jq &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nt"&gt;--arg&lt;/span&gt; w &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$WEATHER&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--arg&lt;/span&gt; now &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$NOW&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s1"&gt;'{
  model: "gemma4:e4b",
  stream: false,
  messages: [
    {role: "system", content: "Suggest one 30 to 45 minute window to go outside today. It must start after the current time and end before sunset. Use only the weather data you are given. If no window fits, set possible to false."},
    {role: "user", content: ("Current time: " + $now + "\nWeather: " + $w)}
  ],
  format: {
    type: "object",
    properties: {possible: {type: "boolean"}, start: {type: "string"}, end: {type: "string"}, reason: {type: "string"}},
    required: ["possible", "start", "end", "reason"]
  }
}'&lt;/span&gt; | curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; http://localhost:11434/api/chat &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; @- | jq &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'.message.content | fromjson'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;format&lt;/code&gt; field constrains the model to that JSON shape, and the answer comes back as a JSON string in &lt;code&gt;.message.content&lt;/code&gt;, which is why the example parses it with &lt;code&gt;fromjson&lt;/code&gt;. In your own code, check the result against the data before you use it: the window must start after now, end before sunset, and fall in hours with a low chance of rain. After sunset there is no valid window, so expect &lt;code&gt;possible: false&lt;/code&gt;. Better still, let plain code pick the window and use the model only for the &lt;code&gt;reason&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to fit it into one week
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Monday or Tuesday: pick the person and the place.&lt;/strong&gt; Write one sentence: "My on-call friend wants to go for a walk without worrying about pages." That sentence is your scope.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Midweek: build the ugly version.&lt;/strong&gt; One input, one output, from the command line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Saturday: take it outside.&lt;/strong&gt; Use it on a real walk, hike or garden. What breaks out there is the best paragraph in your post.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sunday: write.&lt;/strong&gt; Writing has the most weight. Leave yourself at least three hours.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What to put in the write-up
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Who it is for&lt;/strong&gt; and what got them outside, in their words.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why open, concretely.&lt;/strong&gt; "It works in airplane mode on the trail" is stronger than "open source is important".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happened outside.&lt;/strong&gt; The organisers ask for this as a bonus, and it is also the most interesting part to read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What did not work:&lt;/strong&gt; a model that guessed a species with too much confidence, a route that ignored a river. Judges trust posts that show the hard parts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A demo:&lt;/strong&gt; photos from outside, a short video, or a terminal recording.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your agent session,&lt;/strong&gt; if you used one. The organisers suggest DevRelay. It is optional, but it helps judges see how you built it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Next week
&lt;/h2&gt;

&lt;p&gt;A new challenge starts every Monday in October, each with a new theme. We post DevOps project ideas for each one on the day it launches.&lt;/p&gt;

&lt;p&gt;And if you want to practise classic open source pull requests this month (they do not count for the DEV Challenge, but they are still worth learning), our &lt;a href="https://devops-daily.com/hacktoberfest" rel="noopener noreferrer"&gt;7-day Hacktoberfest challenge&lt;/a&gt; gives you one small, reviewed PR a day.&lt;/p&gt;

&lt;p&gt;What will you build to get outside this week? Tell us in the comments, and then go for a walk.&lt;/p&gt;

</description>
      <category>hacktoberfest</category>
      <category>ai</category>
      <category>devops</category>
      <category>opensource</category>
    </item>
    <item>
      <title>GITHUB_TOKEN Is Now 377 Characters Long. We Tested Which Redaction Rules Still Catch It</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/githubtoken-is-now-377-characters-long-we-tested-which-redaction-rules-still-catch-it-3e4e</link>
      <guid>https://dev.to/devopsdaily/githubtoken-is-now-377-characters-long-we-tested-which-redaction-rules-still-catch-it-3e4e</guid>
      <description>&lt;p&gt;On October 2, GitHub &lt;a href="https://github.blog/changelog/2026-10-02-stateless-github-app-installation-tokens-rolled-out/" rel="noopener noreferrer"&gt;finished the rollout&lt;/a&gt; of a new format for GitHub App installation tokens. By default, every newly minted &lt;code&gt;ghs_&lt;/code&gt; token, including the &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; that Actions gives each job, is now &lt;code&gt;ghs_&amp;lt;app id&amp;gt;_&amp;lt;JWT&amp;gt;&lt;/code&gt; instead of a 40-character opaque string. GitHub's checklist tells you to look at length checks, database columns, proxies, and "logging and secret redaction rules that only match the legacy token pattern." We wanted to know how many of those rules actually break, so we ran 40 real tokens through the usual regexes, and one of them through three popular secret scanners and two databases.&lt;/p&gt;

&lt;p&gt;Most of them missed. The latest gitleaks release, with its default rules, found nothing. The common hand-written patterns hid only the first 40 to 46 characters, and in our sample that part never changed from one token to the next, so the "redacted" log line still held everything an attacker needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Actions token we measured is 377 characters,&lt;/strong&gt; not the "about 520" in GitHub's changelog. All 40 samples had the same length and shape: &lt;code&gt;ghs_15368_&lt;/code&gt; plus a three-part JWT signed with ES256.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The first 46 characters were identical in all 40 tokens.&lt;/strong&gt; They are the app ID of GitHub Actions and a base64url JWT header that never changed. A rule that redacts only that part hides nothing secret.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Five common patterns redacted the whole token in 0 of 40 cases.&lt;/strong&gt; Three never matched at all, so the full token would stay in the log. Two matched only the 40 or 46 public characters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub's own first recommended regex left the hyphen out.&lt;/strong&gt; The May 15 version, &lt;code&gt;ghs_[A-Za-z0-9\._]{36,}&lt;/code&gt;, fully matched 8 of 40 tokens, because 32 had a &lt;code&gt;-&lt;/code&gt; in the JWT. GitHub corrected it on May 26, and the fixed regex matched all 40.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;gitleaks 8.30.1 (latest release, March 21) reported 0 findings.&lt;/strong&gt; trufflehog 3.97.9 found a 46-character fragment and marked it unverified. detect-secrets 1.5.0 flagged the line. Fixes for gitleaks and trufflehog have been open pull requests since July.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage fails loudly, except one case.&lt;/strong&gt; Postgres and default MySQL 8.4 reject the token in a &lt;code&gt;VARCHAR(255)&lt;/code&gt; column. MySQL with strict mode off stores the first 255 characters with only a warning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two custom rules, gitleaks and trufflehog, catch the full 377-character token today.&lt;/strong&gt; Both are below and in the companion repo.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Code, logs, or a database that handles GitHub App installation tokens or the Actions &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/BurntSushi/ripgrep" rel="noopener noreferrer"&gt;ripgrep&lt;/a&gt; (&lt;code&gt;rg&lt;/code&gt;) to search a codebase for old patterns&lt;/li&gt;
&lt;li&gt;Optional: a GitHub account to fork the test repo and run it against your own rules&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What changed, and when
&lt;/h2&gt;

&lt;p&gt;GitHub announced the change on &lt;a href="https://github.blog/changelog/2026-04-24-notice-about-upcoming-new-format-for-github-app-installation-tokens/" rel="noopener noreferrer"&gt;April 24&lt;/a&gt;. The staged rollout started on April 27 with the Actions &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; and first-party integrations, then moved to all App installation tokens from mid-May to late June. The October 2 post marks it as complete.&lt;/p&gt;

&lt;p&gt;What changed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Format:&lt;/strong&gt; &lt;code&gt;ghs_APPID_JWT&lt;/code&gt;. The &lt;code&gt;ghs_&lt;/code&gt; prefix stays.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Length:&lt;/strong&gt;"~520 characters" in GitHub's words, and it "will vary based on the data stored within it."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contents:&lt;/strong&gt; GitHub says the JWT "contains details about the token such as the target installation, the application, and basic validation details," and that clients must not validate it or depend on its contents.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What did not change: permissions, repository scoping, the one-hour expiration, and the REST endpoint that mints the token. GitHub Enterprise Server is not affected.&lt;/p&gt;

&lt;p&gt;One date still matters. GitHub added a temporary &lt;code&gt;X-GitHub-Stateless-S2S-Token&lt;/code&gt; request header in &lt;a href="https://github.blog/changelog/2026-05-15-github-app-installation-tokens-per-request-override-header/" rel="noopener noreferrer"&gt;May&lt;/a&gt; so apps could force either format. Apps that send &lt;code&gt;disabled&lt;/code&gt; to keep getting old tokens lose that option on &lt;strong&gt;November 30, 2026&lt;/strong&gt; , when GitHub stops respecting the header.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we tested without leaking a token
&lt;/h2&gt;

&lt;p&gt;We did not have a third-party GitHub App to mint tokens from, so we used the token every Actions job already has. GitHub's April notice says the new format covers "GitHub App installation server-to-server tokens, including Actions &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;." A workflow in a private test repository read &lt;code&gt;secrets.GITHUB_TOKEN&lt;/code&gt; and passed it to small Python scripts that print only structure: lengths, separator positions, the decoded JWT header, the claim names, and how many characters each regex leaves visible. No step printed the payload values or the signature, and we checked the downloaded logs for any token value before making the repository public.&lt;/p&gt;

&lt;p&gt;Two workflows ran on October 5, 2026, on &lt;code&gt;ubuntu-latest&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;check.yml&lt;/code&gt;: one token through the shape script, the regex list, gitleaks, trufflehog, detect-secrets, and Postgres and MySQL service containers.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sample.yml&lt;/code&gt;: 40 matrix jobs, one token each, to see whether the results hold across tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the shape script's output from the recorded check run:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;check.yml, Token shape step&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# TOKEN is the job's GITHUB_TOKEN; the script prints structure only
$ python3 scripts/shape.py
length: 377
prefix: ghs_
separators (index, char): [(3, '_'), (9, '_'), (46, '.'), (290, '.'), (321, '-'), (360, '-')]
app id segment: 15368 (public: the app's numeric id)
segment lengths header/payload/signature: 36 243 86
chars used outside [A-Za-z0-9]: ['-']
header: {"alg":"ES256","typ":"JWT"}
payload claim names: ['aud', 'ctx', 'exp', 'iat', 'iss', 'jti', 'ver']
exp - iat (s): 3600
signature bytes: 64

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Put together, a current &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ghs_15368_eyJhbGciOiJFUzI1NiIsInR5cCI6IkpXVCJ9.&amp;lt;243-character payload&amp;gt;.&amp;lt;86-character signature&amp;gt;
| ____________________________________________ |
  46 characters, the same in all 40 samples:
  ghs_ + 15368 (the app ID of GitHub Actions) + _ +
  the base64url of {"alg":"ES256","typ":"JWT"}

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few things we did not expect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;377, not 520.&lt;/strong&gt; All 40 Actions tokens were exactly 377 characters. GitHub's "about 520" probably describes tokens for other apps, which carry more claims. We could not mint one, so we cannot confirm that. Plan for at least 520, as GitHub asks, and do not assume a fixed length either way.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The prefix is public.&lt;/strong&gt; The 46-character head (&lt;code&gt;ghs_15368_&lt;/code&gt; plus the header segment) had one SHA-256 value across all 40 samples. It decodes to &lt;code&gt;{"alg":"ES256","typ":"JWT"}&lt;/code&gt;, which anyone can encode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The JWT uses base64url, so &lt;code&gt;-&lt;/code&gt; and &lt;code&gt;_&lt;/code&gt; appear inside it.&lt;/strong&gt; 32 of the 40 tokens had a &lt;code&gt;-&lt;/code&gt;, 27 had a &lt;code&gt;_&lt;/code&gt;, and only 3 had neither. This detail breaks the most regexes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Actions masks the token in its own logs, but nobody else does.&lt;/strong&gt; GitHub masks secret values in workflow logs and says that this &lt;a href="https://docs.github.com/en/actions/reference/security/secure-use#use-secrets-for-sensitive-information" rel="noopener noreferrer"&gt;redaction is not guaranteed&lt;/a&gt;. Your application logs, proxies, error trackers, and log pipelines do not get that masking at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The redaction test
&lt;/h2&gt;

&lt;p&gt;For each pattern, the scripts ran &lt;code&gt;re.search&lt;/code&gt; against the token and counted the characters outside the first match. Every pattern here starts with &lt;code&gt;ghs_&lt;/code&gt;, which appears once per token, so that is also what a &lt;code&gt;re.sub&lt;/code&gt; replacement would leave visible. The patterns are ones you find in real redaction code: GitHub's old documented shape, the regexes from the current gitleaks and trufflehog releases, the two regexes GitHub recommended in May, and the regexes from the open scanner fixes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Matched&lt;/th&gt;
&lt;th&gt;Whole token redacted (40 tokens)&lt;/th&gt;
&lt;th&gt;Most characters left visible&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ghs_[0-9a-zA-Z]{36}&lt;/code&gt; (legacy shape)&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;377&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;`(?:ghu&lt;/td&gt;
&lt;td&gt;ghs)_[0-9a-zA-Z]{36}` (gitleaks 8.30.1)&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;\bghs_[A-Za-z0-9_]{36}\b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;377&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;`(?:ghu&lt;/td&gt;
&lt;td&gt;ghs)&lt;em&gt;[A-Za-z0-9&lt;/em&gt;]{36}`&lt;/td&gt;
&lt;td&gt;first 40 characters&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;trufflehog 3.97.9 GitHub detector&lt;/td&gt;
&lt;td&gt;first 46 characters&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;331&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ghs_[A-Za-z0-9\._]{36,}&lt;/code&gt; (GitHub, May 15)&lt;/td&gt;
&lt;td&gt;up to the first &lt;code&gt;-&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;86&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ghs_[A-Za-z0-9\.\-_]{36,}&lt;/code&gt; (GitHub, May 26)&lt;/td&gt;
&lt;td&gt;whole token&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gitleaks PR #2193&lt;/td&gt;
&lt;td&gt;whole token&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;trufflehog PR #5156&lt;/td&gt;
&lt;td&gt;whole token&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Tokens fully redacted, out of 40 real GITHUB_TOKENs&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Legacy ghs_ + 36&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Old shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gitleaks 8.30.1 rule&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Old shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Word-bounded 36&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Old shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Underscore 36&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Old shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;trufflehog 3.97.9 rule&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Old shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub, May 15&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;New-format aware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub, May 26&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;New-format aware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gitleaks PR #2193&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;New-format aware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;trufflehog PR #5156&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;New-format aware&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Each pattern run with Python re against 40 Actions tokens minted on 2026-10-05 (sample.yml in The-DevOps-Daily/ghs-token-check). 'Fully redacted' means re.sub would replace all 377 characters.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The failures fall into three groups.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never matches, so the full token stays in the log.&lt;/strong&gt; The legacy shape and the gitleaks rule expect 36 letters or digits right after &lt;code&gt;ghs_&lt;/code&gt;. The new token has five digits (&lt;code&gt;15368&lt;/code&gt;) and then an underscore, so the match fails at character 10. The word-bounded variant allows underscores, but its 36 characters end in the middle of the header, where no word boundary exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Matches the public prefix only.&lt;/strong&gt; Allow underscores and the pattern matches &lt;code&gt;ghs_15368_&lt;/code&gt; plus 30 header characters. Make it open-ended (&lt;code&gt;{36,255}&lt;/code&gt;, which is what trufflehog uses) and it runs to the first &lt;code&gt;.&lt;/code&gt; and stops at 46 characters. Either way, the payload and the full signature stay in the log line. In our sample the part that disappeared was the same in every token, so anyone who reads the redacted line can paste the prefix back. The result is a working token for as long as it lives, which for an installation token is up to an hour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stops at the first hyphen.&lt;/strong&gt; GitHub's May 15 changelog recommended &lt;code&gt;ghs_[A-Za-z0-9\._]{36,}&lt;/code&gt; "to match both new and current format tokens." A &lt;a href="http://web.archive.org/web/20260519035919/https://github.blog/changelog/2026-05-15-github-app-installation-tokens-per-request-override-header/" rel="noopener noreferrer"&gt;snapshot from May 19&lt;/a&gt; shows that version. An editor's note dated May 26 updated it to &lt;code&gt;ghs_[A-Za-z0-9\.\-_]{36,}&lt;/code&gt;. The first version has no &lt;code&gt;-&lt;/code&gt;, and base64url uses &lt;code&gt;-&lt;/code&gt;. In our sample, the match stopped somewhere in the signature on 32 of 40 tokens and left 8 to 86 characters visible. For redaction this is the least bad failure, because the hidden part always includes the payload, and in most of those 32 tokens part of the signature too. In one sample the whole signature stayed visible. For &lt;strong&gt;validation&lt;/strong&gt; it is worse: code that checks a token with &lt;code&gt;fullmatch&lt;/code&gt; against the May 15 regex rejects every token that contains a hyphen. In our sample that was 80% of real tokens.&lt;/p&gt;

&lt;p&gt;If you copied GitHub's regex in the second half of May, check which version you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the scanners found
&lt;/h2&gt;

&lt;p&gt;The check workflow wrote the token into a file as &lt;code&gt;GITHUB_TOKEN=&amp;lt;token&amp;gt;&lt;/code&gt; and pointed the latest release of each scanner at it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;gitleaks v8.30.1
gitleaks findings: 0
trufflehog 3.97.9
trufflehog findings: 1
  Github verified False raw_len 46
detect-secrets 1.5.0
detect-secrets findings: 2
  JSON Web Token
  GitHub Token

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;gitleaks 8.30.1 found nothing.&lt;/strong&gt; Its &lt;code&gt;github-app-token&lt;/code&gt; rule is &lt;code&gt;(?:ghu|ghs)_[0-9a-zA-Z]{36}&lt;/code&gt;, the first "never matches" case above. Someone reported the format change in &lt;a href="https://github.com/gitleaks/gitleaks/issues/2192" rel="noopener noreferrer"&gt;issue #2192&lt;/a&gt; on July 15, and &lt;a href="https://github.com/gitleaks/gitleaks/pull/2193" rel="noopener noreferrer"&gt;PR #2193&lt;/a&gt; with a fix has been open since July 16. The latest release is from March 21. If gitleaks 8.30.1 with its default rules guards your pre-commit hook or CI, its rule cannot match this token shape, and it did not flag our test file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;trufflehog 3.97.9 found a fragment.&lt;/strong&gt; Its detector matched the 46-character prefix and reported that fragment as an unverified finding. Most examples in trufflehog's README run with &lt;code&gt;--results=verified&lt;/code&gt;, and that filter drops this finding. &lt;a href="https://github.com/trufflesecurity/trufflehog/pull/5156" rel="noopener noreferrer"&gt;PR #5156&lt;/a&gt;, open since July 27, adds a pattern for the new format.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;detect-secrets 1.5.0 flagged the line twice,&lt;/strong&gt; once as a JSON Web Token and once as a GitHub Token. Its GitHub pattern also matches only a prefix, but a scanner only needs to flag the line, and it did.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point explains why the same partial match is fine in one tool and bad in another. A &lt;strong&gt;scanner&lt;/strong&gt; that matches 46 characters still points a person at the right line. A &lt;strong&gt;redactor&lt;/strong&gt; that matches 46 characters prints the rest of the token.&lt;/p&gt;

&lt;h2&gt;
  
  
  Storing the token
&lt;/h2&gt;

&lt;p&gt;GitHub asks for columns that "accept at least 520 characters." We inserted the 377-character token into narrow columns to see how each database reacts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;postgres t40:
ERROR: value too long for type character varying(40)
postgres t255:
ERROR: value too long for type character varying(255)
mysql 8.4 default sql_mode: ONLY_FULL_GROUP_BY,STRICT_TRANS_TABLES,NO_ZERO_IN_DATE,NO_ZERO_DATE,ERROR_FOR_DIVISION_BY_ZERO,NO_ENGINE_SUBSTITUTION
mysql strict t255:
ERROR 1406 (22001) at line 1: Data too long for column 'tok' at row 1
mysql non-strict t255:
stored_len
255

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Postgres and MySQL 8.4 with its default strict mode fail on insert, which is the good outcome: the error is in your logs the first time a new token arrives. MySQL with &lt;code&gt;sql_mode&lt;/code&gt; cleared, which some older applications set on purpose, stores the first 255 characters and keeps going. The insert succeeds with a warning that is easy to miss, and the failure shows up later as a 401 from the GitHub API.&lt;/p&gt;

&lt;p&gt;Installation tokens last an hour, so most apps cache them rather than store them for long. Look anyway: caches with fixed-size keys or values, cookie-backed sessions, and old migrations from when the token was always 40 characters.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to change
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Find the old patterns.&lt;/strong&gt; Search code and config for anything that encodes the old shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# regexes that expect the old 36-character body, and anything mentioning ghs_&lt;/span&gt;
rg &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nt"&gt;--hidden&lt;/span&gt; &lt;span class="nt"&gt;--glob&lt;/span&gt; &lt;span class="s1"&gt;'!.git'&lt;/span&gt; &lt;span class="nt"&gt;--glob&lt;/span&gt; &lt;span class="s1"&gt;'!node_modules'&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'ghs_'&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'gh\[[a-z]+\]_'&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'\{36\}'&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;

&lt;span class="c"&gt;# columns that might hold a token and are too short&lt;/span&gt;
rg &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'varchar\((40|64|100|128|255)\)'&lt;/span&gt; &lt;span class="nt"&gt;--glob&lt;/span&gt; &lt;span class="s1"&gt;'*.sql'&lt;/span&gt; &lt;span class="nt"&gt;--glob&lt;/span&gt; &lt;span class="s1"&gt;'*migration*'&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your log pipeline or error tracker has its own scrubbing rules, check those too. A rule written for the old shape lives wherever someone pasted it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Redact with a pattern that covers the whole JWT.&lt;/strong&gt; GitHub's corrected regex works:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;

&lt;span class="c1"&gt;# GitHub's recommendation as of May 26, 2026; matches old and new ghs_ tokens
&lt;/span&gt;&lt;span class="n"&gt;GHS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ghs_[A-Za-z0-9\.\-_]{36,}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;redact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;GHS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ghs_[REDACTED]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is greedy, so it also eats a trailing &lt;code&gt;.&lt;/code&gt; or &lt;code&gt;-&lt;/code&gt; that ends a sentence. For redaction that is the right tradeoff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Do not validate the structure.&lt;/strong&gt; GitHub says clients "must not take a dependency on the contents of this JWT." If you check tokens at all, check the &lt;code&gt;ghs_&lt;/code&gt; prefix and a generous maximum length, and let the GitHub API decide if the token is valid.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Patch your scanners until the upstream fixes ship.&lt;/strong&gt; Both of these caught the full 377-character token in the recorded run (&lt;code&gt;gitleaks with config/gitleaks.toml findings: 1, secret_len 377&lt;/code&gt; and &lt;code&gt;CustomRegex verified False raw_len 377&lt;/code&gt;):&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Custom rules for the new installation token format&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;gitleaks&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="c"&gt;# .gitleaks.toml: keep the default rules, add one&lt;/span&gt;
&lt;span class="nn"&gt;[extend]&lt;/span&gt;
&lt;span class="py"&gt;useDefault&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

&lt;span class="nn"&gt;[[rules]]&lt;/span&gt;
&lt;span class="py"&gt;id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"github-app-installation-token-jwt"&lt;/span&gt;
&lt;span class="py"&gt;description&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"GitHub App installation token, stateless ghs_&amp;lt;app id&amp;gt;_&amp;lt;JWT&amp;gt; format"&lt;/span&gt;
&lt;span class="py"&gt;regex&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;'''ghs_[0-9]+_eyJ[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+'''&lt;/span&gt;
&lt;span class="py"&gt;keywords&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"ghs_"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;trufflehog&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# trufflehog --config trufflehog.yaml&lt;/span&gt;
&lt;span class="na"&gt;detectors&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;GitHubInstallationTokenJWT&lt;/span&gt;
    &lt;span class="na"&gt;keywords&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;ghs_&lt;/span&gt;
    &lt;span class="na"&gt;regex&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;token&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ghs_[0-9]+_eyJ[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+'&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A trufflehog custom detector without a verification endpoint reports its findings as unverified. A run with &lt;code&gt;--results=verified&lt;/code&gt; drops them too, so let unverified results through for this detector.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Size storage for the documented length, not the one you measured.&lt;/strong&gt; Use &lt;code&gt;text&lt;/code&gt; in Postgres, or at least &lt;code&gt;VARCHAR(1024)&lt;/code&gt; where you need a bound, and make sure MySQL runs in strict mode so a short column fails on insert.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Remove the override header.&lt;/strong&gt; If your app sends &lt;code&gt;X-GitHub-Stateless-S2S-Token: disabled&lt;/code&gt; to keep old-format tokens, that stops working on November 30, 2026. Fix whatever needed it before then.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test your own rules
&lt;/h2&gt;

&lt;p&gt;The workflows, scripts, and recorded results are public. Fork the repository, add your redaction or validation regex to &lt;code&gt;patterns.txt&lt;/code&gt;, and run the workflows. The test uses the fork's own &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;, prints no token, and reports how much of a live token your rule leaves visible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/The-DevOps-Daily/ghs-token-check" rel="noopener noreferrer"&gt;The-DevOps-Daily/ghs-token-check on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What we could not test
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tokens from other GitHub Apps.&lt;/strong&gt; We measured only the Actions &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;. GitHub says other installation tokens are about 520 characters, and their payloads may differ. Every regex that covers the full JWT structure should still match, but we did not run one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub's own secret scanning and push protection.&lt;/strong&gt; Testing them means pushing a live token to a repository, which we chose not to do, and it would say little about your own pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commercial log scrubbers.&lt;/strong&gt; We tested regexes and three open source scanners, not the built-in rules in Datadog, Splunk, or Sentry. Run your own provider's rule through the test repo if you can export it as a regex.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How long the prefix stays fixed.&lt;/strong&gt; In our 40 samples the header segment never changed. GitHub could change the signing algorithm or add a header field at any time, so treat "the first 46 characters are public" as a property of today's tokens, not a guarantee.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;GitHub warned about this change for five months and lists "logging and secret redaction rules" in its own checklist. The tools many teams use for that job have not caught up: the latest gitleaks release does not detect the new tokens, trufflehog sees only a fragment that it cannot verify, and the common hand-written patterns redact a prefix that, in our tests, was identical in every token. The fixes are small. Use a pattern that covers the whole JWT, add one custom rule to each scanner, give the column room for 520 characters or more, and test the result against a real token instead of a 40-character example in a unit test.&lt;/p&gt;

&lt;p&gt;If you are reviewing the rest of your GitHub Actions setup, our posts on &lt;a href="https://dev.to/devopsdaily/github-started-enforcing-self-hosted-runner-versions-we-tested-what-happens-3nm8"&gt;self-hosted runner version enforcement&lt;/a&gt; and on &lt;a href="https://dev.to/devopsdaily/you-cannot-rotate-a-secret-you-cannot-find-13n4"&gt;finding secrets before you have to rotate them&lt;/a&gt; cover two related changes.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/github-app-installation-tokens-redaction" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>githubactions</category>
      <category>githubapps</category>
      <category>secretsmanagement</category>
    </item>
    <item>
      <title>Knowing What Your AI Feature Costs Before Finance Does</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/knowing-what-your-ai-feature-costs-before-finance-does-303e</link>
      <guid>https://dev.to/devopsdaily/knowing-what-your-ai-feature-costs-before-finance-does-303e</guid>
      <description>&lt;p&gt;The invoice for your model provider arrives with one line per model. Finance reads it and asks a fair question: which feature spent that? If three features share one API key, the invoice cannot tell you, and neither can the provider dashboard. You find out the hard way, when one of them grows.&lt;/p&gt;

&lt;p&gt;We wanted to know how far off the obvious guesses are, so we measured. Three features on one key, two models on &lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;DigitalOcean Serverless Inference&lt;/a&gt;, 151 recorded calls, and every number below comes from the &lt;code&gt;usage&lt;/code&gt; block the provider returned. The feature we expected to be expensive was not the problem. The one-word classifier was.&lt;/p&gt;

&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;On Kimi K2.6, an incident answer the engineer read was about 100 tokens long. The request billed about 2,900 tokens, and 928 of them were reasoning tokens that nobody reads.&lt;/li&gt;
&lt;li&gt;A one-word alert label used a median of 566.5 reasoning tokens on Kimi K2.6 and 184.5 on DeepSeek V4.1 Flash, with single calls as high as 1,590 and 1,783. &lt;strong&gt;On mean costs, one label cost between 32 and 48 percent of a full incident investigation.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Turning reasoning off for the label (&lt;code&gt;reasoning_effort: "none"&lt;/code&gt; on Kimi K2.6) cut the cost per 1,000 labels 41 times at list price. It also changed 6 of the 20 labels.&lt;/li&gt;
&lt;li&gt;The agent loop was not the multiplier we expected. In 19 of 20 runs the model asked for all four tools in one step, so an investigation took two calls, not the ten we had feared.&lt;/li&gt;
&lt;li&gt;The same prompt used anywhere from 551 to 1,652 reasoning tokens across ten runs. Per-request cost is a distribution, not a number.&lt;/li&gt;
&lt;li&gt;The fix is one wrapper: every call carries the feature and tenant that caused it, and writes the provider's usage in OpenTelemetry GenAI attribute names.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Node.js 20 or later&lt;/li&gt;
&lt;li&gt;An API key for an OpenAI-compatible endpoint that returns a &lt;code&gt;usage&lt;/code&gt; block (we used DigitalOcean Serverless Inference)&lt;/li&gt;
&lt;li&gt;A rough idea of per-token pricing: you pay one rate for input tokens and a higher rate for output tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The scripts, the fixtures, every raw response's usage and the report are in the companion repo:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/The-DevOps-Daily/ai-feature-cost-ledger" rel="noopener noreferrer"&gt;The-DevOps-Daily/ai-feature-cost-ledger on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The invoice has the wrong shape
&lt;/h2&gt;

&lt;p&gt;A cloud bill has the same problem, and the cloud answer is &lt;a href="https://devops-daily.com/posts/cloud-cost-allocation-tags-aws-gcp-azure" rel="noopener noreferrer"&gt;cost allocation tags&lt;/a&gt;. Model spend has no tags. The provider sees an API key and a model name, so that is what it bills by. Everything you would want to group by, such as the feature, the customer and the environment, exists only in your code at the moment you make the call.&lt;/p&gt;

&lt;p&gt;There is a second problem, and it is the one that surprised us. Even per request, what a user sees has little to do with what you pay for. A request bills four things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input tokens&lt;/strong&gt; : the system prompt, the conversation so far, tool definitions and tool results, every time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cached input tokens&lt;/strong&gt; : the part of the input the provider served from its prompt cache, at a lower rate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output tokens&lt;/strong&gt; : the answer&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning tokens&lt;/strong&gt; : thinking the model does before it answers, billed as output and usually never shown&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first and last are the ones people underestimate. So we built three small features that share one key and measured each.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three features
&lt;/h2&gt;

&lt;p&gt;All three run on the same key, which is the point: the provider sees one stream of calls.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Incident assistant.&lt;/strong&gt; An engineer asks why &lt;code&gt;checkout-api&lt;/code&gt; is returning 502s. The assistant has four tools: list pods, read logs, list recent deploys and read metrics. The tools return fixed text describing pods that are OOM-killed after a deploy raised a cache size, so every run sees the same incident. We ran it two ways: as a tool-calling agent, and single-shot with all four tool outputs pasted into one prompt. The system prompt asks for at most three short sentences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ticket summary.&lt;/strong&gt; Two sentences for an account manager from a nine-message support thread.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alert classifier.&lt;/strong&gt; One word, &lt;code&gt;page&lt;/code&gt;, &lt;code&gt;ticket&lt;/code&gt; or &lt;code&gt;ignore&lt;/code&gt;, for 20 different alert lines.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each incident condition ran 10 times on each of two models, DeepSeek V4.1 Flash and Kimi K2.6. The classifier ran all 20 alerts under four settings. Both models report reasoning tokens and cached tokens separately in &lt;code&gt;usage&lt;/code&gt;, which is why we picked them. All 40 incident answers identified the intended root cause (a keyword check for the memory exhaustion and the cache change behind it), so the cost comparisons below are between answers that found the right cause.&lt;/p&gt;

&lt;p&gt;Prices are DigitalOcean's published Standard rates per million tokens, read from the &lt;a href="https://docs.digitalocean.com/products/inference/details/pricing/" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt; on 5 October 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$0.006&lt;/td&gt;
&lt;td&gt;$1.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6&lt;/td&gt;
&lt;td&gt;$0.95&lt;/td&gt;
&lt;td&gt;$0.19&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What one answer actually billed
&lt;/h2&gt;

&lt;p&gt;Here is the single-shot incident answer on Kimi K2.6, straight from the report file:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ai-feature-cost-ledger&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# one incident answer on Kimi K2.6, median of 10 runs&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;jq &lt;span class="s1"&gt;'.incident["kimi-k2.6 single"] | {medianInput, medianOutput, medianReasoning, medianVisibleAnswer, medianBilledPerVisible}'&lt;/span&gt; data/report.json
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"medianInput"&lt;/span&gt;: 1893,
  &lt;span class="s2"&gt;"medianOutput"&lt;/span&gt;: 1023.5,
  &lt;span class="s2"&gt;"medianReasoning"&lt;/span&gt;: 927.5,
  &lt;span class="s2"&gt;"medianVisibleAnswer"&lt;/span&gt;: 102,
  &lt;span class="s2"&gt;"medianBilledPerVisible"&lt;/span&gt;: 30.6
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The engineer read a 102-token answer. The request billed a median of 1,893 input tokens and 1,024 output tokens, and 928 of the output tokens were reasoning. For every token the engineer read, the request billed about 30 (the median of the ten per-run ratios). On DeepSeek V4.1 Flash the same question billed about 17 tokens per token read, because it reasoned less: a median of 126 reasoning tokens against Kimi's 928.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One incident answer: what was billed vs what was read&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input (prompt and tool output)&lt;/td&gt;
&lt;td&gt;1893 tokens&lt;/td&gt;
&lt;td&gt;Billed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning (billed as output)&lt;/td&gt;
&lt;td&gt;927.5 tokens&lt;/td&gt;
&lt;td&gt;Billed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer the engineer read&lt;/td&gt;
&lt;td&gt;102 tokens&lt;/td&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Kimi K2.6, single-shot, median of 10 runs. Each bar is its own median, so the bars need not add up. Source: data/report.json in the companion repo.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Reasoning tokens are in &lt;code&gt;completion_tokens&lt;/code&gt;, and you pay the output rate for them. If you estimate cost by counting the tokens in the answer text, you miss most of the output bill on a reasoning model. Read &lt;code&gt;usage.completion_tokens_details.reasoning_tokens&lt;/code&gt; where the provider reports it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The agent loop was not the problem
&lt;/h2&gt;

&lt;p&gt;We expected the agent to be the expensive version. An agent resends the whole conversation on every step: system prompt, tool definitions, every earlier tool result. Ten sequential steps means paying for the first prompt ten times.&lt;/p&gt;

&lt;p&gt;That is not what happened. In 19 of the 20 agent runs, the model asked for all four tools in its first step, received the results, and answered in the second. The one exception, a Kimi run, asked for three tools, then the fourth, then answered: three calls. With two calls, the resend costs one extra copy of a short prompt.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Single-shot, median&lt;/th&gt;
&lt;th&gt;Agent, median&lt;/th&gt;
&lt;th&gt;Agent / single-shot&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash&lt;/td&gt;
&lt;td&gt;$0.000926&lt;/td&gt;
&lt;td&gt;$0.001518&lt;/td&gt;
&lt;td&gt;1.64x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6&lt;/td&gt;
&lt;td&gt;$0.005892&lt;/td&gt;
&lt;td&gt;$0.005871&lt;/td&gt;
&lt;td&gt;1.00x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those are list prices with no cache discount, which is the fair comparison here (the cache section explains why). On Kimi the agent was no dearer at all, because it reasoned less once the tool results were in front of it.&lt;/p&gt;

&lt;p&gt;This does not mean agent loops are cheap. It means the step count decides, and the step count belongs to the model and the task, not to you. A model that calls one tool per step, or a task where each tool result decides the next call, turns two calls into eight. Measure steps per request as its own number, and alert on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-word answer that cost the most
&lt;/h2&gt;

&lt;p&gt;The classifier is the feature nobody worries about. The prompt is roughly 50 to 90 tokens and the answer is one word. At list price, it cost $0.48 per 1,000 labels on DeepSeek V4.1 Flash with default settings, and $2.45 on Kimi K2.6.&lt;/p&gt;

&lt;p&gt;Almost all of that is reasoning. The answer was 2 to 4 tokens every time. The reasoning was not:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reasoning tokens spent on a one-word alert label&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;th&gt;Samples&lt;/th&gt;
&lt;th&gt;Min&lt;/th&gt;
&lt;th&gt;Median&lt;/th&gt;
&lt;th&gt;p95&lt;/th&gt;
&lt;th&gt;Max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash, default&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;52 tokens&lt;/td&gt;
&lt;td&gt;184.5 tokens&lt;/td&gt;
&lt;td&gt;1218 tokens&lt;/td&gt;
&lt;td&gt;1783 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash, effort=low&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;46 tokens&lt;/td&gt;
&lt;td&gt;107.5 tokens&lt;/td&gt;
&lt;td&gt;832 tokens&lt;/td&gt;
&lt;td&gt;1138 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6, default&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;154 tokens&lt;/td&gt;
&lt;td&gt;566.5 tokens&lt;/td&gt;
&lt;td&gt;1168 tokens&lt;/td&gt;
&lt;td&gt;1590 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6, effort=none&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;20 alerts per setting, one call each. The visible answer was 2 to 4 tokens every time. Source: data/feature-runs.jsonl.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The tail is where the money goes. A DNS failure rate of 14 percent for three minutes drew 1,783 reasoning tokens from DeepSeek before it said &lt;code&gt;ticket&lt;/code&gt;. A node memory-pressure alert drew 1,218. On DeepSeek at default settings, 14 of the 20 alerts used between 50 and 250. Nothing in the request tells you in advance which call will be the expensive one.&lt;/p&gt;

&lt;p&gt;Put next to the incident assistant, this is the number that changes priorities. Comparing mean costs, &lt;strong&gt;one label cost 32 to 48 percent of a full agent investigation&lt;/strong&gt; : 32 percent on DeepSeek at list price, 48 percent on Kimi with the cache discounts both features got. A feature that runs on every alert is billed at a third to a half of a feature that runs when an engineer is paged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning reasoning down, and what it changes
&lt;/h2&gt;

&lt;p&gt;Both models accept a &lt;code&gt;reasoning_effort&lt;/code&gt; parameter on this endpoint, but not the same values, and the values do not mean the same thing. DeepSeek V4.1 Flash rejected &lt;code&gt;none&lt;/code&gt; and &lt;code&gt;minimal&lt;/code&gt; (&lt;code&gt;reasoning_effort must be one of [low high xhigh max] for this model&lt;/code&gt;) and accepted &lt;code&gt;low&lt;/code&gt;. Kimi K2.6 accepted all three, but on the probe alert &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;minimal&lt;/code&gt; used more reasoning than the default (1,210 and 1,021 tokens against 656), and only &lt;code&gt;none&lt;/code&gt; turned it off. That is one call each, recorded in &lt;code&gt;data/effort-probe.json&lt;/code&gt;, so read it as "check what the setting does", not as a rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ai-feature-cost-ledger&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# the alert classifier on Kimi K2.6, default vs reasoning_effort none&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;jq &lt;span class="s1"&gt;'.classifier["kimi default"] | {runs, medianOutput, medianReasoning, maxOutput, medianVisible, costPer1000}'&lt;/span&gt; data/report.json
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"runs"&lt;/span&gt;: 20,
  &lt;span class="s2"&gt;"medianOutput"&lt;/span&gt;: 569.5,
  &lt;span class="s2"&gt;"medianReasoning"&lt;/span&gt;: 566.5,
  &lt;span class="s2"&gt;"maxOutput"&lt;/span&gt;: 1593,
  &lt;span class="s2"&gt;"medianVisible"&lt;/span&gt;: 3,
  &lt;span class="s2"&gt;"costPer1000"&lt;/span&gt;: 2.4195
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;jq &lt;span class="s1"&gt;'.classifier["kimi effort=none"] | {runs, medianOutput, medianReasoning, costPer1000, agreesWithDefault}'&lt;/span&gt; data/report.json
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"runs"&lt;/span&gt;: 20,
  &lt;span class="s2"&gt;"medianOutput"&lt;/span&gt;: 2,
  &lt;span class="s2"&gt;"medianReasoning"&lt;/span&gt;: 0,
  &lt;span class="s2"&gt;"costPer1000"&lt;/span&gt;: 0.0381,
  &lt;span class="s2"&gt;"agreesWithDefault"&lt;/span&gt;: 14
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Cost per 1,000 alert labels&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6, default&lt;/td&gt;
&lt;td&gt;2.4451$&lt;/td&gt;
&lt;td&gt;Kimi K2.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6, effort=none&lt;/td&gt;
&lt;td&gt;0.06$&lt;/td&gt;
&lt;td&gt;Kimi K2.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash, default&lt;/td&gt;
&lt;td&gt;0.4812$&lt;/td&gt;
&lt;td&gt;DeepSeek V4.1 Flash&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash, effort=low&lt;/td&gt;
&lt;td&gt;0.2671$&lt;/td&gt;
&lt;td&gt;DeepSeek V4.1 Flash&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Mean cost of 20 real calls per setting, scaled to 1,000, at DigitalOcean list prices on 5 Oct 2026 with no cache discount (costPer1000NoCache). Source: data/report.json.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;With reasoning off, Kimi produced 2 output tokens per label. At list price, the cost per 1,000 labels fell from $2.45 to $0.06, 41 times less. As billed, with the cache discounts some of those calls got, it fell from $2.42 to $0.04, 63 times less. DeepSeek at &lt;code&gt;low&lt;/code&gt; saved 1.8 times either way, and its tail did not go away: one call still used 1,138 reasoning tokens.&lt;/p&gt;

&lt;p&gt;The catch is in the last field. Without reasoning, Kimi gave a different label on 6 of the 20 alerts. Four moved up to &lt;code&gt;page&lt;/code&gt;: the disk at 91 percent, a node under memory pressure, a CI queue delay and a single OOM-killed pod. The other two moved from &lt;code&gt;ignore&lt;/code&gt; to &lt;code&gt;ticket&lt;/code&gt;. DeepSeek at &lt;code&gt;low&lt;/code&gt; changed 2 labels. Each alert ran once per setting, so these are observed differences, not a measured rate, and we have no ground truth for the alerts, so we cannot say which setting was right. We can say that "turn reasoning off" is a product decision about who gets woken up, not a free saving. Run your own labelled set through both settings before you flip it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching is a discount you cannot schedule
&lt;/h2&gt;

&lt;p&gt;Prompt caching is the other big lever. DeepSeek V4.1 Flash bills cached input at $0.006 per million tokens against $0.30 uncached, 50 times less. On the nine single-shot runs that hit the cache, the median cost was $0.000342, against $0.000925 for the same tokens at list price.&lt;/p&gt;

&lt;p&gt;Two things stop you from budgeting on that number:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The first call always paid full price.&lt;/strong&gt; The first call of every incident condition had zero cached tokens, and so did the first call of the second Kimi agent run, even though the run before it had just sent the same prefix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What repeats in our test does not repeat in production.&lt;/strong&gt; We sent the identical incident ten times, so later runs found almost the entire single-shot prompt in cache (1,984 of 2,019 tokens on DeepSeek). A real incident brings new logs and new metrics. Only the shared prefix (system prompt and tool definitions, 576 tokens on DeepSeek here) can repeat across different incidents, which is why the agent comparison above uses list prices.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Record &lt;code&gt;cache_read&lt;/code&gt; tokens per call and watch the hit rate as its own metric. Put the stable part of every prompt first and the per-request part last, so the prefix the cache can match is as long as possible. Then budget at list price and treat the cache as a discount, not a plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tag every call with the feature that caused it
&lt;/h2&gt;

&lt;p&gt;None of the numbers above helps finance unless each call says which feature it belongs to. The fix is boring, and that is the point. Every model call goes through one function. The function knows the feature and the tenant, reads the &lt;code&gt;usage&lt;/code&gt; block from the response, prices it, and writes one row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/ledger.mjs, trimmed&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;appendFileSync&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node:fs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;chat&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./client.mjs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;splitUsage&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./cost.mjs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;meteredChat&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;feature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tenant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ledger&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;splitUsage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toISOString&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;app.feature&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;feature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// who caused this call&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;app.tenant&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;tenant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// who you would bill or rate limit&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.operation.name&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;chat&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.provider.name&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;digitalocean&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.request.model&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.response.model&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.usage.input_tokens&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// includes cached tokens&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.usage.cache_read.input_tokens&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.usage.output_tokens&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// includes reasoning tokens&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gen_ai.usage.reasoning.output_tokens&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reasoning&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;app.cost_usd&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="nf"&gt;appendFileSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the pricing function it calls, which is where most homemade cost trackers go wrong:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/cost.mjs, trimmed&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;splitUsage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;u&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt_tokens&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt_tokens_details&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;cached_tokens&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cache_read_input_tokens&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completion_tokens&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reasoning&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completion_tokens_details&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;reasoning_tokens&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;uncached&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;reasoning&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pricing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;models&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt; &lt;span class="c1"&gt;// USD per million tokens, with source and date&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;splitUsage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cachedRate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cachedInput&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;uncached&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;cachedRate&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;e6&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three details matter more than the rest.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use the provider's numbers, not your own estimate.&lt;/strong&gt; The &lt;code&gt;usage&lt;/code&gt; block is what you are billed for. A tokenizer on your side counts the text you sent; it does not see reasoning tokens or what the cache served.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split before you price.&lt;/strong&gt; &lt;code&gt;prompt_tokens&lt;/code&gt; already includes cached tokens and &lt;code&gt;completion_tokens&lt;/code&gt; already includes reasoning tokens. Add them on top and you bill yourself twice for the same tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the fields the way OpenTelemetry does.&lt;/strong&gt; The &lt;a href="https://github.com/open-telemetry/semantic-conventions-genai" rel="noopener noreferrer"&gt;GenAI semantic conventions&lt;/a&gt; define &lt;code&gt;gen_ai.usage.input_tokens&lt;/code&gt;, &lt;code&gt;gen_ai.usage.output_tokens&lt;/code&gt;, &lt;code&gt;gen_ai.usage.cache_read.input_tokens&lt;/code&gt; and &lt;code&gt;gen_ai.usage.reasoning.output_tokens&lt;/code&gt;, and state that the cached and reasoning counts are included in the totals, the same rule as above. The conventions are still marked Development, so pin the version you emit. The names still mean a JSON line today can become a span attribute later without a rename.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once every row carries &lt;code&gt;app.feature&lt;/code&gt;, the roll-up is a &lt;code&gt;GROUP BY&lt;/code&gt;. The 151 calls in this experiment cost $0.19 as billed and $0.22 at list price. As billed, the 80 classifier calls cost more than the 41 agent calls; at list price, the agent calls were slightly ahead:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Calls&lt;/th&gt;
&lt;th&gt;As billed&lt;/th&gt;
&lt;th&gt;List price&lt;/th&gt;
&lt;th&gt;Reasoning tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;alert-classifier&lt;/td&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;td&gt;$0.0641&lt;/td&gt;
&lt;td&gt;$0.0651&lt;/td&gt;
&lt;td&gt;23,500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;incident-assistant (agent)&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;$0.0619&lt;/td&gt;
&lt;td&gt;$0.0713&lt;/td&gt;
&lt;td&gt;7,948&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;incident-assistant (single-shot)&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;$0.0566&lt;/td&gt;
&lt;td&gt;$0.0733&lt;/td&gt;
&lt;td&gt;11,831&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ticket-summary&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;$0.0053&lt;/td&gt;
&lt;td&gt;$0.0066&lt;/td&gt;
&lt;td&gt;2,958&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The mix of calls here is our test plan, not a production workload, so the shares mean nothing on their own. The shape is the lesson: the feature with the shortest answers spent the most reasoning tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a gateway fits
&lt;/h2&gt;

&lt;p&gt;You can stop at the wrapper. It is about 30 lines, it lives in your code, and it adds no one who sees your prompts beyond the inference provider. Keeping it consistent gets harder when the calls come from many services in several languages, and it reports spend without enforcing a budget. That is the problem the AI gateway and LLM observability vendors sell into, and they solve the tagging part in similar ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://portkey.ai/docs/product/observability/metadata" rel="noopener noreferrer"&gt;Portkey&lt;/a&gt; takes an &lt;code&gt;x-portkey-metadata&lt;/code&gt; header with any keys you choose, such as &lt;code&gt;feature&lt;/code&gt; and &lt;code&gt;_user&lt;/code&gt;, and has budget limits in front of the provider.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.helicone.ai/features/advanced-usage/custom-properties" rel="noopener noreferrer"&gt;Helicone&lt;/a&gt; uses one header per property, &lt;code&gt;Helicone-Property-&amp;lt;Name&amp;gt;&lt;/code&gt;, and lets you segment requests and cost by them.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://langfuse.com/docs/observability/features/token-and-cost-tracking" rel="noopener noreferrer"&gt;Langfuse&lt;/a&gt; records usage and cost per generation. It infers cost from model definitions it ships for OpenAI, Anthropic and Google models, so for a model like DeepSeek V4.1 Flash on DigitalOcean you add your own definition with the prices, or send usage and cost yourself. Ingested values take priority over inferred ones, and you should send them.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.braintrust.dev/docs/observe" rel="noopener noreferrer"&gt;Braintrust&lt;/a&gt; logs traces with token counts, cost and metadata alongside its evaluations, which fits the question this post leaves open: whether a cheaper reasoning setting still labels alerts correctly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade-offs are the usual ones. A proxy gateway adds a network hop to every model call, and a hosted one is another company that sees your prompts. Helicone, Langfuse and Portkey's gateway all publish source you can run yourself, which answers the second point and makes the first your problem. Whichever you pick, check that it records the provider's reported usage, including reasoning and cached tokens, rather than estimating from the text. Counting only what the user saw would have put the incident assistant's token bill far too low in this test: the median billed-to-read ratio per condition ran from 17 to 40, and single investigations from 14 to 50.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we could not conclude
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Whether failed attempts are billed.&lt;/strong&gt; The 151 recorded calls all succeeded on the first attempt, so the client's retries for 429 and 5xx never fired in the record. A request that fails or times out on the client side may still reach the provider and be billed, and nothing we recorded can show whether it was.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Which classifier labels were right.&lt;/strong&gt; We have no ground truth for the 20 alerts, so the 6 changed labels are a difference, not an error rate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether reasoning helps the incident answers.&lt;/strong&gt; All 40 identified the intended root cause by our keyword check, which is not a full accuracy review, so we cannot price the accuracy reasoning buys on this task. A harder incident might change that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything about other models.&lt;/strong&gt; Two models, one endpoint, one day. Others reason more or less, and the OpenAI GPT-5 models on the same pricing page were not available on our account tier, so we could not include them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steady-state cache hit rates.&lt;/strong&gt; Our repeated prompts overstate hits for anything that changes per request.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Price every request from the provider's &lt;code&gt;usage&lt;/code&gt; block. The answer text is not the bill.&lt;/li&gt;
&lt;li&gt;On reasoning models, reasoning tokens can be most of the output bill, even for a one-word answer. Track &lt;code&gt;reasoning_tokens&lt;/code&gt; as its own metric and look at the tail, not the median.&lt;/li&gt;
&lt;li&gt;Steps per request decides what an agent costs. The median was 2 here: 19 investigations took two calls and one took three. Measure it and alert when it climbs.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;reasoning_effort&lt;/code&gt; is a large lever and a product decision. Test the labels it changes before you ship it.&lt;/li&gt;
&lt;li&gt;Budget at list price. Treat the prompt cache as a discount, put the stable prefix first, and watch the hit rate.&lt;/li&gt;
&lt;li&gt;Wrap every model call once, tag it with the feature and tenant, and use the OpenTelemetry GenAI names. Then finance gets one line per feature, and you get it before they ask.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For two other ways a model bill grows quietly, see &lt;a href="https://dev.to/devopsdaily/jev-and-the-classification-problem-hiding-in-your-llm-bill-191j"&gt;why most LLM calls are classification&lt;/a&gt; and &lt;a href="https://dev.to/devopsdaily/your-semantic-cache-answers-the-question-next-door-3d55"&gt;what a semantic cache really saves&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/ai-feature-cost-before-finance-does" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>finops</category>
      <category>ai</category>
      <category>llm</category>
      <category>opentelemetry</category>
    </item>
    <item>
      <title>A Repo per Agent: What Cloudflare Artifacts Changes About Git</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Sat, 03 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/a-repo-per-agent-what-cloudflare-artifacts-changes-about-git-3m5b</link>
      <guid>https://dev.to/devopsdaily/a-repo-per-agent-what-cloudflare-artifacts-changes-about-git-3m5b</guid>
      <description>&lt;p&gt;Most Git workflows assume that changes come from people, arrive at human speed, and get reviewed by other people. Coding agents break all three assumptions. A team that runs fifty agents in parallel does not have fifty developers; it has fifty processes that may all want to land work in the same minute. When they all push to one branch, they spend much of that time fetching, rebasing and retrying.&lt;/p&gt;

&lt;p&gt;On October 1, 2026, Cloudflare announced new capabilities for &lt;strong&gt;Artifacts&lt;/strong&gt; , its Git-compatible store that you drive from Cloudflare Workers, together with a contest to build what comes next on top of it. Artifacts is in open beta on the Workers Paid plan. Its core idea is easy to say and has big consequences: stop sharing one repository between workers, and give each unit of work its own. This post explains what Artifacts is, measures why a shared branch stops working as agents multiply, shows the repo-per-agent pattern in code, and works out what it costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Artifacts is a Git-compatible repository store you create and control from a Worker. Any Git client can clone and push with a short-lived bearer token.&lt;/li&gt;
&lt;li&gt;Cloudflare's own guidance is one repo per unit of autonomous work: 10,000 agents, 10,000 repos. Repos are cheap to create and fork.&lt;/li&gt;
&lt;li&gt;In our local experiment, failed pushes grew roughly with the square of the number of agents sharing one branch: 40 agents produced about 700 failed push attempts.&lt;/li&gt;
&lt;li&gt;A repo per agent removes the collisions but moves the hard part to merging the work back. Cloudflare left that layer open and is running a contest for it.&lt;/li&gt;
&lt;li&gt;Pricing: 10,000 operations and 1 GB free per month, then $0.15 per 1,000 operations and $0.50 per GB-month. Each repo is capped at 1 GB.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;To follow the code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Cloudflare account on the &lt;strong&gt;Workers Paid&lt;/strong&gt; plan (Artifacts is not available on the free plan)&lt;/li&gt;
&lt;li&gt;Wrangler 4.145.0 or later, so &lt;code&gt;wrangler types&lt;/code&gt; knows the Artifacts binding&lt;/li&gt;
&lt;li&gt;Git 2.x on the machine that clones and pushes&lt;/li&gt;
&lt;li&gt;Basic familiarity with Workers bindings and &lt;code&gt;wrangler.toml&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The push experiment below needs Git, Bash and standard Unix tools (&lt;code&gt;bc&lt;/code&gt;, &lt;code&gt;paste&lt;/code&gt;, &lt;code&gt;sort&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  What Artifacts is
&lt;/h2&gt;

&lt;p&gt;An Artifacts &lt;strong&gt;namespace&lt;/strong&gt; holds repositories. Each &lt;strong&gt;repository&lt;/strong&gt; is a real Git remote: you get an HTTPS URL of the form &lt;code&gt;https://&amp;lt;ACCOUNT_ID&amp;gt;.artifacts.cloudflare.net/git/&amp;lt;namespace&amp;gt;/&amp;lt;repo&amp;gt;.git&lt;/code&gt;, and any Git client can talk to it. Authentication is a bearer token passed as an extra HTTP header, so the token never lands in &lt;code&gt;.git/config&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Clone and push with a repo-scoped token; nothing is written to the remote URL&lt;/span&gt;
git &lt;span class="nt"&gt;-c&lt;/span&gt; http.extraHeader&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$ARTIFACTS_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; clone &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ARTIFACTS_REMOTE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; work
&lt;span class="nb"&gt;cd &lt;/span&gt;work
&lt;span class="c"&gt;# ...edit and commit...&lt;/span&gt;
git &lt;span class="nt"&gt;-c&lt;/span&gt; http.extraHeader&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$ARTIFACTS_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; push origin main

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What makes it different from a hosted Git server is the control plane. A Worker gets a binding with methods to create, import, fork, list and delete repositories, mint read or write tokens with a lifetime in seconds, and read history and files without cloning:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Need&lt;/th&gt;
&lt;th&gt;Binding call&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;New empty repo&lt;/td&gt;
&lt;td&gt;&lt;code&gt;env.ARTIFACTS.create(name)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Copy from GitHub or another remote&lt;/td&gt;
&lt;td&gt;&lt;code&gt;env.ARTIFACTS.import({ source: { url }, target: { name } })&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Branch off a reviewed baseline&lt;/td&gt;
&lt;td&gt;&lt;code&gt;repo.fork(name, { defaultBranchOnly: true })&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credentials for one agent&lt;/td&gt;
&lt;td&gt;&lt;code&gt;repo.createToken("write", 900)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inspect what an agent did&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;repo.log({ ref: "main" })&lt;/code&gt;, &lt;code&gt;repo.readFile({ ref, path })&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clean up&lt;/td&gt;
&lt;td&gt;&lt;code&gt;env.ARTIFACTS.delete(name)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few more pieces make it a platform rather than storage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Workers Builds&lt;/strong&gt; can deploy a Worker from an Artifacts repo. Pushes to &lt;code&gt;main&lt;/code&gt; run the deploy command, and once you enable builds for preview branches, pushes to other branches produce preview URLs. Only &lt;code&gt;main&lt;/code&gt; can be the production branch for now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repository events&lt;/strong&gt; (creates, forks, pushes, clones and so on) can be delivered to Workers Queues, so a push can start a test run or a review agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Jurisdictions&lt;/strong&gt; keep a namespace's data in the US or the EU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics&lt;/strong&gt; per repository: operations, pulls, pushes and error rates.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Under the hood, Cloudflare describes each repo as a single logical instance that it can route to from any region, with data replicated synchronously across data centers and copied to object storage in the background. It has not published how concurrent pushes to the same ref are ordered, which matters for the next section.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why one shared branch falls apart
&lt;/h2&gt;

&lt;p&gt;Git protects a branch with a compare-and-swap. A push says "move &lt;code&gt;main&lt;/code&gt; from commit A to commit B". If someone else moved &lt;code&gt;main&lt;/code&gt; first, the server refuses, because B was built on a history that is no longer the tip. Two agents are enough to see it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;two agents, one branch&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# both agents cloned the same baseline and made one commit each&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git push origin main &lt;span class="c"&gt;# in agent-a&lt;/span&gt;
To ../shared.git
   35bfd50..11799a8 main -&amp;gt; main
&lt;span class="nv"&gt;$ &lt;/span&gt;git push origin main &lt;span class="c"&gt;# in agent-b&lt;/span&gt;
To ../shared.git
 &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;rejected] main -&amp;gt; main &lt;span class="o"&gt;(&lt;/span&gt;fetch first&lt;span class="o"&gt;)&lt;/span&gt;
error: failed to push some refs to &lt;span class="s1"&gt;'../shared.git'&lt;/span&gt;
hint: Updates were rejected because the remote contains work that you &lt;span class="k"&gt;do
&lt;/span&gt;hint: not have locally. This is usually caused by another repository pushing
hint: to the same ref. You may want to first integrate the remote changes
hint: &lt;span class="o"&gt;(&lt;/span&gt;e.g., &lt;span class="s1"&gt;'git pull ...'&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; before pushing again.
hint: See the &lt;span class="s1"&gt;'Note about fast-forwards'&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s1"&gt;'git push --help'&lt;/span&gt; &lt;span class="k"&gt;for &lt;/span&gt;details.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Agent B now has to fetch, rebase and try again. That is fine for two humans. To see what happens with more writers, we ran a small experiment: N agents each commit one file (no two agents touch the same file, so the work never conflicts), are launched concurrently, and push to &lt;code&gt;main&lt;/code&gt; of one shared repository, retrying immediately with &lt;code&gt;git pull --rebase&lt;/code&gt; until the push lands.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="c"&gt;# N agents each commit one file and push to the same branch of one shared repo,&lt;/span&gt;
&lt;span class="c"&gt;# retrying with fetch + rebase until the push lands. Prints total attempts.&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;
&lt;span class="nv"&gt;N&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;W&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$W&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
git init &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;--bare&lt;/span&gt; &lt;span class="nt"&gt;-b&lt;/span&gt; main shared.git
git clone &lt;span class="nt"&gt;-q&lt;/span&gt; shared.git seed 2&amp;gt;/dev/null
&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;seed &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git &lt;span class="nt"&gt;-c&lt;/span&gt; user.email&lt;span class="o"&gt;=&lt;/span&gt;s@x &lt;span class="nt"&gt;-c&lt;/span&gt; user.name&lt;span class="o"&gt;=&lt;/span&gt;seed commit &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; baseline &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git push &lt;span class="nt"&gt;-q&lt;/span&gt; origin main&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq &lt;/span&gt;1 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$N&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;git clone &lt;span class="nt"&gt;-q&lt;/span&gt; shared.git &lt;span class="s2"&gt;"a&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
&lt;/span&gt;agent&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="s2"&gt;"a&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"agent-&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;.txt"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; git add .&lt;span class="p"&gt;;&lt;/span&gt; git &lt;span class="nt"&gt;-c&lt;/span&gt; user.email&lt;span class="o"&gt;=&lt;/span&gt;a@x &lt;span class="nt"&gt;-c&lt;/span&gt; user.name&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"agent-&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; commit &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"agent-&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  &lt;span class="nv"&gt;tries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0
  &lt;span class="k"&gt;until &lt;/span&gt;git push &lt;span class="nt"&gt;-q&lt;/span&gt; origin main 2&amp;gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    &lt;/span&gt;&lt;span class="nv"&gt;tries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;tries+1&lt;span class="k"&gt;))&lt;/span&gt;
    git &lt;span class="nt"&gt;-c&lt;/span&gt; user.email&lt;span class="o"&gt;=&lt;/span&gt;a@x &lt;span class="nt"&gt;-c&lt;/span&gt; user.name&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"agent-&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; pull &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;--rebase&lt;/span&gt; origin main 2&amp;gt;/dev/null
  &lt;span class="k"&gt;done
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$tries&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; ../fails-&lt;span class="nv"&gt;$1&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq &lt;/span&gt;1 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$N&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;agent &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &amp;amp; &lt;span class="k"&gt;done&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;wait
&lt;/span&gt;&lt;span class="nv"&gt;total&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;fails-&lt;span class="k"&gt;*&lt;/span&gt; | &lt;span class="nb"&gt;paste&lt;/span&gt; &lt;span class="nt"&gt;-sd&lt;/span&gt;+ | bc&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;max&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;fails-&lt;span class="k"&gt;*&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;commits&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;shared.git rev-list &lt;span class="nt"&gt;--count&lt;/span&gt; main&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"agents=&lt;/span&gt;&lt;span class="nv"&gt;$N&lt;/span&gt;&lt;span class="s2"&gt; failed_pushes=&lt;/span&gt;&lt;span class="nv"&gt;$total&lt;/span&gt;&lt;span class="s2"&gt; worst_agent_retries=&lt;/span&gt;&lt;span class="nv"&gt;$max&lt;/span&gt;&lt;span class="s2"&gt; commits_on_main=&lt;/span&gt;&lt;span class="nv"&gt;$commits&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$W&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We ran it three times for each size on a 4-core Raspberry Pi 4 with Git 2.39 and a bare repository on local disk. The counter is failed push attempts; the script does not record why each one failed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;race.sh, three runs per size&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;n &lt;span class="k"&gt;in &lt;/span&gt;5 10 20 40&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do for &lt;/span&gt;run &lt;span class="k"&gt;in &lt;/span&gt;1 2 3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt; ./race.sh &lt;span class="nv"&gt;$n&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
&lt;/span&gt;&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;5 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;4 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;6
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;5 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;4 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;6
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;5 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;4 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;6
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;45 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;9 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;11
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;45 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;9 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;11
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;45 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;9 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;11
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;190 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;19 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;21
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;186 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;19 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;21
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;190 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;19 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;21
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;40 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;711 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;36 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;41
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;40 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;716 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;37 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;41
&lt;span class="nv"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;40 &lt;span class="nv"&gt;failed_pushes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;702 &lt;span class="nv"&gt;worst_agent_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;38 &lt;span class="nv"&gt;commits_on_main&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;41

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;commits_on_main&lt;/code&gt; is the baseline plus one commit per agent, so every agent's work landed in the end. The cost is in the retries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failed push attempts when N agents share one branch&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;5 agents&lt;/th&gt;
&lt;th&gt;10 agents&lt;/th&gt;
&lt;th&gt;20 agents&lt;/th&gt;
&lt;th&gt;40 agents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Measured (median of 3 runs)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;190&lt;/td&gt;
&lt;td&gt;711&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;N(N-1)/2&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;190&lt;/td&gt;
&lt;td&gt;780&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;race.sh on a Raspberry Pi 4, Git 2.39, local bare repository. No agents edited the same file, so every rebase succeeded; real conflicts make it worse.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A simple model explains the shape. If the agents move in rounds and only one push wins each round, every other pending push fails, and the total is N(N-1)/2. Our results sit close to that model. They fall below it at 40 agents, which is what you expect when an agent fetches several new commits in one rebase and skips some rounds. The model gives 499,500 failed pushes at 1,000 agents; we did not measure that scale.&lt;/p&gt;

&lt;p&gt;Keep the limits of this test in mind. It is a conflict-free workload on local disk with immediate retries; backoff and spread-out arrival times would lower the numbers, and a hosted server adds a network round trip to every attempt. Real agents editing the same code would also hit merge conflicts that a retry loop cannot fix. And the work around Git has its own budgets: GitHub generally allows up to 80 content-generating requests per minute and 500 per hour across its API and web interface, and REST and GraphQL share a limit of 100 concurrent requests. Those limits apply to things like creating branches, pull requests and comments, not to the Git pushes counted here.&lt;/p&gt;

&lt;p&gt;The usual answer is to give each agent its own branch in the shared repo. That removes the race on &lt;code&gt;main&lt;/code&gt;, but the repository is still one shared object for permissions, clone size and blast radius. A token that can push to the repo can usually push to every unprotected branch in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repo-per-agent pattern
&lt;/h2&gt;

&lt;p&gt;Artifacts makes the next step cheap: a fork per session. A reviewed &lt;strong&gt;baseline&lt;/strong&gt; repo holds the code you trust. Every agent session gets its own fork, a write token that only works on that fork and expires in minutes, and nothing else.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;baseline&lt;/strong&gt; reviewed main&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;agent-a-s41&lt;/strong&gt; fork + 15 min token&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;agent-b-s42&lt;/strong&gt; fork + 15 min token&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;agent-c-s43&lt;/strong&gt; fork + 15 min token&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orchestrator&lt;/strong&gt; tests, review, merge&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Connections:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;baseline -&amp;gt; agent-a-s41 (fork)&lt;/li&gt;
&lt;li&gt;baseline -&amp;gt; agent-b-s42 (fork)&lt;/li&gt;
&lt;li&gt;baseline -&amp;gt; agent-c-s43 (fork)&lt;/li&gt;
&lt;li&gt;agent-a-s41 -&amp;gt; Orchestrator (push event)&lt;/li&gt;
&lt;li&gt;agent-b-s42 -&amp;gt; Orchestrator (push event)&lt;/li&gt;
&lt;li&gt;agent-c-s43 -&amp;gt; Orchestrator (push event)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Worker that hands out sessions is short. Configure the binding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"agent-sessions"&lt;/span&gt;
&lt;span class="py"&gt;main&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"src/index.ts"&lt;/span&gt;
&lt;span class="py"&gt;compatibility_date&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"2026-10-02"&lt;/span&gt;

&lt;span class="nn"&gt;[[artifacts]]&lt;/span&gt;
&lt;span class="py"&gt;binding&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"ARTIFACTS"&lt;/span&gt;
&lt;span class="py"&gt;namespace&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"agents"&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then fork the baseline per session and return a scoped, short-lived token. This assumes a repo named &lt;code&gt;baseline&lt;/code&gt; already exists in the namespace with your reviewed code on &lt;code&gt;main&lt;/code&gt; (create it with &lt;code&gt;create&lt;/code&gt; and push, or &lt;code&gt;import&lt;/code&gt; it from GitHub), and that you ran &lt;code&gt;wrangler types&lt;/code&gt; so &lt;code&gt;Env&lt;/code&gt; includes the binding and the secret:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/index.ts: one fork and one 15-minute write token per agent session&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="na"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Env&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;method&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST only&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;405&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="c1"&gt;// This endpoint hands out write tokens: only your orchestrator may call it.&lt;/span&gt;
    &lt;span class="c1"&gt;// ORCHESTRATOR_SECRET is a Worker secret (wrangler secret put ORCHESTRATOR_SECRET).&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ORCHESTRATOR_SECRET&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Unauthorized&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;401&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;session&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;session&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="c1"&gt;// Keep names readable but unique: validate the inputs and add a random suffix&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-z0-9&lt;/span&gt;&lt;span class="se"&gt;]{1,24}&lt;/span&gt;&lt;span class="sr"&gt;$/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-z0-9&lt;/span&gt;&lt;span class="se"&gt;]{1,24}&lt;/span&gt;&lt;span class="sr"&gt;$/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;agent and session must be lowercase letters and digits&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;-&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;-&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;crypto&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;randomUUID&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="c1"&gt;// `using` disposes the repo handle when the block ends, as the binding requires&lt;/span&gt;
    &lt;span class="nx"&gt;using&lt;/span&gt; &lt;span class="nx"&gt;baseline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ARTIFACTS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;baseline&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fork&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fork&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;defaultBranchOnly&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="nx"&gt;using&lt;/span&gt; &lt;span class="nx"&gt;repo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ARTIFACTS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;repo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createToken&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;write&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;900&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// 900 s = 15 minutes&lt;/span&gt;

    &lt;span class="c1"&gt;// plaintext is the Git token string; expiresAt tells the agent when to stop&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;remote&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;fork&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;remote&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;plaintext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;expiresAt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;expiresAt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent clones the remote, works, and pushes with the token. When it is done, the orchestrator does not need to clone anything to see what happened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Inspect a finished session without cloning it&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;summarize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Env&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;using&lt;/span&gt; &lt;span class="nx"&gt;repo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ARTIFACTS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;commits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;repo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;main&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt; &lt;span class="c1"&gt;// newest first&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;plan&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;repo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;readFile&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;main&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PLAN.md&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;commits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;plan&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;plan&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;plan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;text&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cloudflare's best-practice notes add three habits worth copying:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Give each token the least it needs, for as little time as possible.&lt;/strong&gt; Read tokens for indexing and review, write tokens only for the agent doing the work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fork from a reviewed baseline&lt;/strong&gt; instead of copying files into each new repo, so every session starts from the same known state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep run metadata out of the tree.&lt;/strong&gt; Attach prompts, model output and run IDs with &lt;code&gt;git notes&lt;/code&gt;, so the commit holds the work and the notes hold the context.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And one of our own: delete forks once their work is merged or rejected. Storage is billed, and a pile of abandoned agent repos is the 2026 version of a pile of abandoned feature branches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting the work back is the hard part
&lt;/h2&gt;

&lt;p&gt;A repo per agent does not make conflicts go away. It moves them from push time, where they cost retries, to integration time, where something has to decide what lands. That is a better place for them, because you can be deliberate there, but someone still has to build it. Cloudflare says so directly: Artifacts is the storage primitive, and its contest asks developers to build the coordination, review and merge layer on top.&lt;/p&gt;

&lt;p&gt;The options teams use today, from simplest to most ambitious:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Orchestrator merges in order.&lt;/strong&gt; One process fetches finished forks, rebases each onto the baseline, runs the tests and pushes. It is a merge queue with exactly one writer to the baseline, so the race from the experiment above cannot happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human review gate.&lt;/strong&gt; Each fork's push event creates a review item. With Workers Builds, a branch push can also produce a preview URL to look at before anything merges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent reviewers.&lt;/strong&gt; A second agent with a read-only token reviews the diff and either approves it into the queue or sends it back. Keep the merge itself with the orchestrator, so no reviewing agent ever holds a write token to the baseline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best-of-N.&lt;/strong&gt; Fork several sessions from the same baseline for the same task, test them all, and merge only the winner. Forks are cheap enough that this stops being wasteful.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are new ideas. What changes is that the per-session repository, its credentials and its cleanup are now an API call instead of something you script around a hosted Git service.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Artifacts bills two things. The numbers below are from the pricing page. It lists October 14, 2026 as the day billing starts; the announcement says October 15:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Included each month&lt;/th&gt;
&lt;th&gt;Then&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Operations (create, push, pull, clone and similar)&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;$0.15 per 1,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;1 GB-month&lt;/td&gt;
&lt;td&gt;$0.50 per GB-month&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;To size it, assume one agent session costs seven operations: a fork, a clone, three pushes, one fetch by the reviewer and a delete. That is an assumption, not a published figure; count your own workflow before you budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Estimated monthly operations cost by agent sessions per day&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100/day&lt;/td&gt;
&lt;td&gt;1.65$&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,000/day&lt;/td&gt;
&lt;td&gt;30$&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10,000/day&lt;/td&gt;
&lt;td&gt;313.5$&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100,000/day&lt;/td&gt;
&lt;td&gt;3148.5$&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Assumes 7 billable operations per session, 30 days, and the 10,000 free operations a month. Excludes storage and other Cloudflare charges.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Storage depends on your repos and how quickly you delete forks. It is billed on the average of each day's peak, so deleting a fork stops it from adding up over the month but does not remove that day's peak. Cloudflare has not documented whether a fork shares objects with its parent or counts its full size, so measure that in the beta before you plan around hundreds of long-lived forks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits and open questions
&lt;/h2&gt;

&lt;p&gt;Know these before you commit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1 GB per repository&lt;/strong&gt; and &lt;strong&gt;32 MB per file or blob&lt;/strong&gt;. A repository above 1 GB, or one with a single file above 32 MB, does not fit. Account storage is 1 TB by default and can be raised.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate limits&lt;/strong&gt; of 2,000 requests per 10 seconds per namespace for the control plane, and 2,000 Git requests per 10 seconds per repository. Split busy workloads across namespaces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open beta.&lt;/strong&gt; The API and limits can still change, and it requires the Workers Paid plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production deploys from &lt;code&gt;main&lt;/code&gt; only&lt;/strong&gt; in the Workers Builds integration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Undocumented so far:&lt;/strong&gt; how concurrent pushes to one ref are ordered, whether forks share storage, and what Artifacts guarantees about publishing events. Queues itself delivers at least once and without ordering guarantees, so make event handlers idempotent. Test the rest before you rely on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The contest runs until October 14, 2026. Entries need a 5 to 10 minute demo video, open source code under MIT, Apache or BSD, and instructions to run it. Up to two members of each of the three winning teams are flown to Cloudflare Connect in San Francisco, and first place also gets $25,000 in Cloudflare credits. Details are on &lt;a href="https://blog.cloudflare.com/next-git-platform-on-cloudflare/" rel="noopener noreferrer"&gt;Cloudflare's announcement&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;Shared branches assume few writers. Our small experiment shows what happens when that assumption fails: with immediate retries, failed pushes grew roughly with the square of the number of agents, before a single real conflict appeared. The fix is not a faster retry loop. It is isolation, and Artifacts makes isolation an API call: fork a reviewed baseline per session, hand the agent a 15-minute token that works on nothing else, and read its work back without cloning.&lt;/p&gt;

&lt;p&gt;What Artifacts does not give you is the decision about what merges. Plan that layer first. A single-writer merge queue plus tests is enough to start, and it is the part where your team's judgment matters most. If you want to start small, move one agent workflow, such as dependency updates or test generation, to forks of a baseline and watch the error rate and the bill for a month. The &lt;a href="https://developers.cloudflare.com/artifacts/" rel="noopener noreferrer"&gt;Artifacts documentation&lt;/a&gt; covers the binding, tokens and Workers Builds setup.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/cloudflare-artifacts-repo-per-agent" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>git</category>
      <category>cloudflare</category>
      <category>aiagents</category>
      <category>cloudflareworkers</category>
    </item>
    <item>
      <title>We Took Down Our Mesh VPN Control Plane. Here Is What Kept Working</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Sat, 03 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/we-took-down-our-mesh-vpn-control-plane-here-is-what-kept-working-d6e</link>
      <guid>https://dev.to/devopsdaily/we-took-down-our-mesh-vpn-control-plane-here-is-what-kept-working-d6e</guid>
      <description>&lt;p&gt;When you replace a VPN concentrator with a mesh VPN such as Tailscale, NetBird or ZeroTier, traffic stops flowing through one box. Devices connect to each other directly wherever the network allows it. But the mesh still has a centre: the &lt;strong&gt;coordination server&lt;/strong&gt; (Tailscale calls it the control plane) that hands out keys, addresses, peer lists and access rules. The old question "what happens when the VPN gateway dies?" becomes "what happens when the control server is unreachable?"&lt;/p&gt;

&lt;p&gt;The vendors' answer is reassuring: the data plane keeps working. We wanted to know what that means in practice, so we ran an open-source control server (&lt;a href="https://github.com/juanfont/headscale" rel="noopener noreferrer"&gt;Headscale&lt;/a&gt;) with real Tailscale clients, stopped it, and probed the network about once a minute for 30 minutes. Existing tunnels survived the whole outage. Joining, revoking and renewing did not, and what happens when a client restarts during the outage is the result most worth adding to your runbook.&lt;/p&gt;

&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;During a control outage&lt;/th&gt;
&lt;th&gt;What we measured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Existing tunnels&lt;/td&gt;
&lt;td&gt;Kept working for the full 30 minutes: 27 of 27 probes in each direction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A new device joins&lt;/td&gt;
&lt;td&gt;Failed (&lt;code&gt;NeedsLogin&lt;/code&gt;); it joined by itself 14 to 22 s after control came back&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;An admin revokes a device&lt;/td&gt;
&lt;td&gt;Impossible: the admin API is the control server&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A device's key expires&lt;/td&gt;
&lt;td&gt;The first probe after the expiry time failed; peers enforce it locally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A client restarts (no netmap cache)&lt;/td&gt;
&lt;td&gt;Came back with no address and stayed off the network until control returned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A client restarts (netmap cache on)&lt;/td&gt;
&lt;td&gt;Reachable again in 15 to 19 s in 11 of 12 timed restarts (one took 140 s); stayed reachable in two 10-minute checks, dropped once in another run&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;With the control server up&lt;/th&gt;
&lt;th&gt;What we measured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Policy change blocks a connection&lt;/td&gt;
&lt;td&gt;The first probe after the change already failed (within 6 s, mostly our probe timeout)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deleting a device&lt;/td&gt;
&lt;td&gt;Same: blocked at the first probe&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A Linux machine with &lt;code&gt;tailscale&lt;/code&gt; and &lt;code&gt;tailscaled&lt;/code&gt; installed (we used 1.102.4)&lt;/li&gt;
&lt;li&gt;Basic familiarity with Tailscale or another mesh VPN&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything runs as a normal user on one host, and it does not touch an existing Tailscale install. The scripts are in the companion repo:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/The-DevOps-Daily/tailnet-control-outage" rel="noopener noreferrer"&gt;The-DevOps-Daily/tailnet-control-outage on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;One Headscale process plays the control server. Five &lt;code&gt;tailscaled&lt;/code&gt; processes play devices, each in &lt;strong&gt;userspace-networking mode&lt;/strong&gt; (no TUN device and no root, so they can share one host). Each node serves &lt;code&gt;hello from &amp;lt;name&amp;gt;&lt;/code&gt; on a local port, and a probe fetches that page through another node's SOCKS5 proxy. A probe only succeeds if traffic really crossed the tailnet.&lt;/p&gt;

&lt;p&gt;Each node has one job:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Node&lt;/th&gt;
&lt;th&gt;Job during the outage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;a, b&lt;/td&gt;
&lt;td&gt;A long-lived pair. Never touched.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;c&lt;/td&gt;
&lt;td&gt;Tries to join.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;d&lt;/td&gt;
&lt;td&gt;Gets restarted.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;e&lt;/td&gt;
&lt;td&gt;Its key expires two minutes in.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;a&lt;/strong&gt; SOCKS5 probe&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WireGuard&lt;/strong&gt; direct path&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;b&lt;/strong&gt; hello from b&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Control plane&lt;/strong&gt; stopped during the test

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Headscale 0.29.4&lt;/strong&gt; keys, peers, policy&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Devices&lt;/strong&gt; tailscaled 1.102.4, userspace mode

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;a, b&lt;/strong&gt; long-lived pair&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;c&lt;/strong&gt; joins mid-outage&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;d&lt;/strong&gt; restarted mid-outage&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;e&lt;/strong&gt; key expires mid-outage&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The outage script stops Headscale and then works through a timeline: probe every node, start c (its join attempt runs from about 25 seconds to 70 seconds), try an admin command, let e's key expire at about two minutes, and restart d at five to six minutes, probing roughly once a minute throughout. Another script then brings the control server back and times the recovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  What kept working: the tunnels you already had
&lt;/h2&gt;

&lt;p&gt;The long-lived pair never noticed. In the 30-minute run, a reached b and b reached a on every probe, 27 probes in each direction, the last one 1,793 seconds after the control server stopped. (In one of the shorter runs, a single probe from b to a failed once and the next one succeeded.) Here is the start and the end of that run, from &lt;code&gt;runs/nocache-30min/02-outage.txt&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;16:34:16 t=0s control server stopped (e's key expires at 2026-10-02T16:36:10Z)
16:34:16 t=0s a-&amp;gt;b ok (hello from b) | b-&amp;gt;a ok (hello from a) | a-&amp;gt;d ok (hello from d) | a-&amp;gt;e ok (hello from e)
16:34:41 t=25s a's own status line: offline
...
17:03:09 t=1723s a-&amp;gt;b ok (hello from b) | b-&amp;gt;a ok (hello from a) | a-&amp;gt;d FAIL | a-&amp;gt;e FAIL
17:04:19 t=1793s a-&amp;gt;b ok (hello from b) | b-&amp;gt;a ok (hello from a) | a-&amp;gt;d FAIL | a-&amp;gt;e FAIL

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This matches what &lt;a href="https://tailscale.com/kb/1091/what-happens-if-the-coordination-server-is-down" rel="noopener noreferrer"&gt;Tailscale documents&lt;/a&gt;: every node keeps its peers, their endpoints and the packet filter, and traffic never goes through the coordination server in the first place. Each probe is a fresh HTTP connection, so this shows that new connections between existing peers keep working; we did not hold one long-lived session open.&lt;/p&gt;

&lt;p&gt;Note the third line. While a was moving traffic, &lt;code&gt;tailscale status&lt;/code&gt; showed a itself as &lt;code&gt;offline&lt;/code&gt;, because a could not reach the control server. If your monitoring alerts on that, it will tell you the mesh is down while it is working. Probe real traffic between real nodes instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stopped working
&lt;/h2&gt;

&lt;h3&gt;
  
  
  New devices cannot join
&lt;/h3&gt;

&lt;p&gt;Node c started with a valid, reusable pre-auth key and never got past &lt;code&gt;NeedsLogin&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;t=71s new node c tries to join: timeout waiting for Tailscale service to enter a Running state; check health with "tailscale status" (state NeedsLogin)

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It kept retrying on its own, and got an address 14 to 22 seconds after the control server came back (three runs). Nobody had to touch it. During the outage, though, a replacement laptop, a new CI runner or an autoscaled node simply cannot get on the network.&lt;/p&gt;

&lt;h3&gt;
  
  
  Nobody can revoke anything
&lt;/h3&gt;

&lt;p&gt;The admin interface is the control server. The &lt;code&gt;headscale nodes list&lt;/code&gt; command we ran to find a node to remove failed with &lt;code&gt;context deadline exceeded&lt;/code&gt;. With Tailscale's hosted service the equivalent is the admin console and API, and they are part of the same control plane.&lt;/p&gt;

&lt;p&gt;This is the security cost of the design. Tailscale's own documentation lists it: during an outage, "existing users cannot have their keys revoked." If you need to cut off a stolen laptop or a departing employee during a control plane incident, the mesh cannot do it for you. The device keeps every peer and every rule it had when the outage started.&lt;/p&gt;

&lt;h3&gt;
  
  
  Expiring keys still expire
&lt;/h3&gt;

&lt;p&gt;We set e's key to expire about two minutes into the outage (114 seconds in the 30-minute run). The first probe after that time failed, and every probe after it. The client does this itself: it marks peers whose key has expired and stops talking to them (&lt;a href="https://github.com/tailscale/tailscale/blob/v1.102.4/ipn/ipnlocal/expiry.go" rel="noopener noreferrer"&gt;&lt;code&gt;ipn/ipnlocal/expiry.go&lt;/code&gt;&lt;/a&gt; in the client source).&lt;/p&gt;

&lt;p&gt;So key expiry is enforced by the peers themselves, without the control server. That is good for security, and it means a control outage can become a data plane outage: every device whose key expires during the outage drops off, and cannot renew until control is back. After recovery, e stayed &lt;code&gt;Logged out&lt;/code&gt; until we cleared its expiry on the server, and then it came back within 1 to 2 seconds. In real life an admin extends or clears the expiry, or the user re-authenticates.&lt;/p&gt;

&lt;p&gt;If your tailnet uses short key expiry for servers, a long control outage will take them off the network one by one. Tailscale lets you disable or temporarily extend expiry per device, and a device that is tagged when it first authenticates has key expiry disabled by default (&lt;a href="https://tailscale.com/kb/1028/key-expiry" rel="noopener noreferrer"&gt;key expiry docs&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  The restart trap
&lt;/h2&gt;

&lt;p&gt;Five minutes into the outage we restarted d's &lt;code&gt;tailscaled&lt;/code&gt;, keeping its state directory, the way an unattended upgrade, a reboot or a crashed process does.&lt;/p&gt;

&lt;p&gt;Without the netmap cache, d came back with no address and no peers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;t=317s restarted d's tailscaled while control is down: state NoState, ip , netmap cache files: 0

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its health check said why (from d's log, in &lt;code&gt;log-excerpts.txt&lt;/code&gt; in the repo):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are logged out. The last login error was: fetch control key: Get "http://127.0.0.1:18080/key?v=142": dial tcp 127.0.0.1:18080: connect: connection refused

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;d stayed off the network for the rest of the outage. The client keeps its keys on disk, but without the cache it does not keep the network map: the list of peers, their addresses and the packet filter only live in memory. Without the control server, a restarted client knows who it is but not who anyone else is. It kept retrying, and once the control server was back it logged in again by itself, within 3 to 4 seconds according to its own log.&lt;/p&gt;

&lt;p&gt;During a control incident, the devices that restart are exactly the ones that disappear. A device that is replaced rather than restarted, such as a new Kubernetes node or a container without persistent state, is in the same position as c: it has never been on the network and cannot join until control is back.&lt;/p&gt;

&lt;h3&gt;
  
  
  Netmap caching fixes it
&lt;/h3&gt;

&lt;p&gt;Tailscale has a fix: &lt;strong&gt;netmap caching&lt;/strong&gt;. With it, each client writes its network map to disk and loads it at startup. Tailscale's &lt;a href="https://tailscale.com/blog/making-tailscale-faster" rel="noopener noreferrer"&gt;September 22 post&lt;/a&gt; says it is a feature flag in the current client and is expected to be on by default from version 1.104, after more testing (mobile clients later). The cache only helps a device that restarts with its state directory intact; a fresh replacement has nothing cached.&lt;/p&gt;

&lt;p&gt;In the 1.102 client we tested, the switch is on the server side. The client only writes the cache when the control server grants it the &lt;code&gt;cache-network-maps&lt;/code&gt; node attribute. The &lt;code&gt;TS_USE_CACHED_NETMAP&lt;/code&gt; environment variable defaults to on and works as an off switch. Headscale 0.29 passes node attributes through from its policy file, so turning it on was one block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"acls"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"accept"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"src"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"lab@"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"dst"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"lab@:*"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"nodeAttrs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"attr"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"cache-network-maps"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;d then kept nine files under its state directory (&lt;code&gt;profile-data/&amp;lt;id&amp;gt;/netmap-cache/&lt;/code&gt;), covering itself, its peers, the user, the packet filter, the DERP map and DNS. With the cache on, the same restart came back like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Start: loaded netmap from disk cache; 3 peers
t=352s restarted d's tailscaled while control is down: state Running, ip 100.64.0.3, netmap cache files: 9

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first line is from d's log (&lt;code&gt;log-excerpts.txt&lt;/code&gt;), the second from the outage script, both from the same 10-minute run.&lt;/p&gt;

&lt;p&gt;d had its address immediately, but peers could not reach it straight away. A restarted client gets a new disco key (the key Tailscale uses for path discovery), and peers normally learn it from the control server. With the cache, the client advertises the new key to its peers directly over the tunnel instead; the logs show &lt;code&gt;sending TSMP disco key advertisement&lt;/code&gt; and the peers receiving it. We timed how long until a could reach d again, with a retrying constantly (a failed attempt times out after 5 seconds, and the next one starts a second later):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Restart (netmap cache on)&lt;/th&gt;
&lt;th&gt;Port&lt;/th&gt;
&lt;th&gt;Outage before the first restart&lt;/th&gt;
&lt;th&gt;a reaches d again after&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3 restarts&lt;/td&gt;
&lt;td&gt;new random port&lt;/td&gt;
&lt;td&gt;under 1 minute&lt;/td&gt;
&lt;td&gt;19 s, 16 s, 15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 restarts&lt;/td&gt;
&lt;td&gt;same fixed port&lt;/td&gt;
&lt;td&gt;under 1 minute&lt;/td&gt;
&lt;td&gt;19 s, 15 s, 15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 restarts&lt;/td&gt;
&lt;td&gt;new random port&lt;/td&gt;
&lt;td&gt;6 minutes&lt;/td&gt;
&lt;td&gt;140 s, 15 s, 17 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 restarts&lt;/td&gt;
&lt;td&gt;same fixed port&lt;/td&gt;
&lt;td&gt;6 minutes&lt;/td&gt;
&lt;td&gt;18 s, 15 s, 16 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Time until a peer reaches a restarted node, control server down&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;th&gt;Samples&lt;/th&gt;
&lt;th&gt;Min&lt;/th&gt;
&lt;th&gt;Median&lt;/th&gt;
&lt;th&gt;p95&lt;/th&gt;
&lt;th&gt;Max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;new random port&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;15s&lt;/td&gt;
&lt;td&gt;16.5s&lt;/td&gt;
&lt;td&gt;140s&lt;/td&gt;
&lt;td&gt;140s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;same fixed port&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;15s&lt;/td&gt;
&lt;td&gt;15.5s&lt;/td&gt;
&lt;td&gt;19s&lt;/td&gt;
&lt;td&gt;19s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;12 restarts with the cache-network-maps attribute, Tailscale 1.102.4 and Headscale 0.29.4. Measured from when the restarted daemon answered, with a new attempt about every 6 s. Without the cache, the node was still unreachable after 300 s.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Eleven of the twelve restarts were reachable again in 15 to 19 seconds. One took 140 seconds, on a clean setup, so slow recoveries do happen. The restarts run one after another, so only the first in each row waited the full outage time; within that limit, a longer outage made no consistent difference, and neither did keeping the same UDP port.&lt;/p&gt;

&lt;p&gt;Does the recovered path last? We restarted d once more and probed every 30 seconds for 10 minutes, once with a fixed port and once with a random port. Both times the first probe, right after the restart, failed, and the next 19, over the following 10 minutes, all reached d. In our 10-minute cache outage run, though, d answered for two minutes after its restart, then failed two probes in a row (at 484 and 559 seconds) and only answered again at 630 seconds, just before control returned. We could not reproduce that, so treat a cached restart as a big improvement, not a guarantee.&lt;/p&gt;

&lt;p&gt;Without the cache, the same test (fixed port) gave up after 300 seconds.&lt;/p&gt;

&lt;p&gt;The cache has a security side. The packet filter is on disk too, so a restarted node enforces the rules it had cached, not newer ones. That is the same trade-off a running node already makes during an outage, now extended across restarts.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the control server comes back
&lt;/h2&gt;

&lt;p&gt;Recovery was fast and needed no help, except for the expired key. From the 10-minute run without the cache:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;16:21:27 control server back
16:21:48 c (tried to join during the outage) gets an address after 21s
16:21:55 a-&amp;gt;c reachable after 7s
16:21:55 a-&amp;gt;b ok (hello from b) | b-&amp;gt;a ok (hello from a)
16:21:55 a-&amp;gt;d (restarted during the outage) reachable after 0s
16:22:00 a-&amp;gt;e (key expired during the outage): FAIL; e's state: Logged out.
16:22:01 admin clears e's expiry: a-&amp;gt;e reachable after 1s

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  With the control server up, access changes are fast
&lt;/h2&gt;

&lt;p&gt;The flip side is how quickly the control plane enforces a change when it is up. We changed the policy so that a may no longer reach b, reloaded Headscale with &lt;code&gt;SIGHUP&lt;/code&gt;, and probed again and again until the result changed. Then we restored it, and then deleted b:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;17:04:59 policy changed to block a-&amp;gt;b: blocked after 6s
17:05:00 policy restored: a-&amp;gt;b works again after 1s
17:05:06 node b deleted: a-&amp;gt;b blocked after 5s

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In both blocking cases the first probe after the change already failed. The 5 to 6 seconds is mostly that probe's own 5-second timeout. So with the control server up, a policy change or a removed device takes effect in seconds. With it down, it does not take effect at all. Your access control is only as live as your control plane.&lt;/p&gt;

&lt;p&gt;This is also where &lt;strong&gt;ACLs as code&lt;/strong&gt; pays off. A policy file in Git, reviewed and applied by CI, is a common way to run both Tailscale and Headscale. During an outage you cannot apply it, but once control is back, the reviewed version goes out in seconds. Git tells you which policy you intended at any point; it does not prove that every device received it, so check from the devices too.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we got wrong first
&lt;/h2&gt;

&lt;p&gt;Three mistakes in our first runs are worth sharing, because they are easy to make in your own health checks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;We restarted half of the long-lived pair.&lt;/strong&gt; The first run used b both as a long-lived peer and as the restart target, so after about six minutes the "do existing tunnels survive?" measurement was gone. We added d as a separate restart target and ran it again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Our recovery script read d's address once, at the start.&lt;/strong&gt; A node restarted without a cache has no address until it logs in again, so every probe went to an empty address, and the script reported that d never recovered. d's own log showed it had logged in again seconds after control returned. If you write tailnet health checks, look up addresses on every attempt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Our first probe accepted any response, and our restart check misread a logged-out client.&lt;/strong&gt; A review of the scripts pointed out that the probe only checked for a non-empty body; it now requires curl to succeed and the page to name the node we meant to reach. The restart helper used &lt;code&gt;tailscale status&lt;/code&gt;, which exits with an error when a node is logged out, so it reported the uncached restart as a failed start instead of measuring it; it now uses &lt;code&gt;tailscale status --json&lt;/code&gt;. Every number in this post comes from runs with the fixed scripts.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What this means if you are replacing a VPN
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Count the control plane in your access path.&lt;/strong&gt; For Tailscale's hosted service, that means their availability. For Headscale, it is one process with one database, so back up the database and treat upgrades like any other change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn on netmap caching&lt;/strong&gt; where your client and control server support it, or plan to upgrade to the release where it is the default. In our tests, without it, a restart during the outage took the device off the network until control returned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not restart clients during a control incident.&lt;/strong&gt; Pause automatic client updates and node rotation until control is back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose key expiry on purpose.&lt;/strong&gt; Expiry limits the damage of a stolen key and also turns a long control outage into lost devices. Decide per device class, and alert before keys expire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a way to contain a device that does not need the control plane&lt;/strong&gt; , such as host firewalls or cloud security groups in front of your most sensitive servers that you can change without the mesh. Revoking sessions at your identity provider helps for applications that check it, but it does not close the network path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor real traffic&lt;/strong&gt; , not the self status. &lt;code&gt;tailscale status&lt;/code&gt; called a working node offline.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we could not test
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Relays.&lt;/strong&gt; Everything ran on one host, and the client status showed direct paths; we did not force relayed connections. If your traffic goes through DERP relays, and especially if you run the relay on the same host as Headscale, an outage may take the relays down as well. We did not measure that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Other products.&lt;/strong&gt; NetBird, ZeroTier, Twingate and Teleport split control and data differently. The questions in this post (joins, revocation, expiry, restarts) are the right ones to ask of any of them, but our numbers only describe Tailscale clients with a Headscale control server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tailscale's hosted control plane.&lt;/strong&gt; It may roll out features like netmap caching differently from Headscale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long outages.&lt;/strong&gt; Our longest was 30 minutes. Anything key-related scales with your expiry settings, not with our test.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;With the control server down for 30 minutes, existing tunnels kept working in both directions on every probe.&lt;/li&gt;
&lt;li&gt;New devices could not join, and nobody could revoke a device. Revocation is the real security cost of the outage.&lt;/li&gt;
&lt;li&gt;Key expiry is enforced by the peers. A key that expires during an outage takes its device off the network until someone re-authenticates it.&lt;/li&gt;
&lt;li&gt;Without the netmap cache, a client restarted during the outage came back with no peers and stayed off the network until control returned. With it (the &lt;code&gt;cache-network-maps&lt;/code&gt; node attribute, expected as the default from Tailscale 1.104), peers reached it again in 15 to 19 seconds in 11 of 12 restarts, and it stayed reachable in two 10-minute checks (one earlier run saw an unexplained two-minute drop).&lt;/li&gt;
&lt;li&gt;With the control server up, a policy change or a deleted device took effect at the first probe, within seconds.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/mesh-vpn-control-plane-outage-what-keeps-working" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>networking</category>
      <category>tailscale</category>
      <category>headscale</category>
      <category>wireguard</category>
    </item>
    <item>
      <title>CVE-2026-80521: The Ubuntu Container Escape Is Not Just an Ubuntu Problem</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/devopsdaily/cve-2026-80521-the-ubuntu-container-escape-is-not-just-an-ubuntu-problem-1kbf</link>
      <guid>https://dev.to/devopsdaily/cve-2026-80521-the-ubuntu-container-escape-is-not-just-an-ubuntu-problem-1kbf</guid>
      <description>&lt;p&gt;On September 22, DepthFirst published a working exploit for &lt;strong&gt;CVE-2026-80521&lt;/strong&gt; : an unprivileged process inside a default Docker or Kubernetes container gets a root shell on the host. Their exploit targets one specific Ubuntu 26.04 kernel, but the bug behind it is much wider. Ten days later, Ubuntu's tracker still lists the main kernels for 24.04 LTS and 26.04 LTS as vulnerable, with no fixed package to install.&lt;/p&gt;

&lt;p&gt;Most coverage calls it "the Ubuntu container escape". That framing is too narrow, and it can make you look in the wrong place. The bug is in upstream Linux code that arrived in 6.10 and was later backported to the 6.1 and 6.6 long-term kernels. Upstream has fixed 6.12, 6.18 and 7.1, but as of today there is no fix in the 6.1 or 6.6 branches at all. Debian 12, for example, is listed as vulnerable.&lt;/p&gt;

&lt;p&gt;This post shows how to tell whether a machine has the affected code (in one command that usually works inside a pod too), where fixes exist, and what to do on nodes that cannot be patched yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Detail&lt;/th&gt;
&lt;th&gt;Info&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CVE&lt;/td&gt;
&lt;td&gt;CVE-2026-80521, CVSS 7.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Class&lt;/td&gt;
&lt;td&gt;Use-after-free in the AF_UNIX socket garbage collector, container escape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reported by&lt;/td&gt;
&lt;td&gt;Kyle Zeng&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public exploit&lt;/td&gt;
&lt;td&gt;September 22, 2026 (DepthFirst), built for an Ubuntu 26.04 LTS kernel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vulnerable code&lt;/td&gt;
&lt;td&gt;Linux 6.10 and later, plus backports in 6.1.141+ and 6.6.93+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fixed upstream&lt;/td&gt;
&lt;td&gt;6.12.111, 6.18.53, 7.1.10, 7.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Not fixed upstream&lt;/td&gt;
&lt;td&gt;6.1 and 6.6 long-term branches (as of October 2)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default seccomp&lt;/td&gt;
&lt;td&gt;Docker's default profile allows the calls the exploit needs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quick check&lt;/td&gt;
&lt;td&gt;`grep -cE ' unix_(add&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Shell access to the hosts you want to check, or {% raw %}&lt;code&gt;kubectl&lt;/code&gt; access to a cluster&lt;/li&gt;
&lt;li&gt;For the cluster check: permission to run &lt;code&gt;kubectl debug node/...&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Your distribution's security tracker for the final answer on a distro kernel&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;Unix domain sockets can carry open file descriptors from one process to another (&lt;code&gt;SCM_RIGHTS&lt;/code&gt;). A socket can even be sent over itself, so the kernel needs a garbage collector to find groups of sockets that only reference each other and free them.&lt;/p&gt;

&lt;p&gt;Linux 6.10 replaced that garbage collector with a new design that tracks sockets as vertices and edges in a graph and groups them into strongly connected components (SCCs). The fix commit describes a race in that code: while one thread sends a socket and another closes the sockets involved, the collector can free a vertex but leave it linked in a cached SCC list. The next garbage collection run walks that list and touches freed memory. That use-after-free is what the exploit turns into a host root shell.&lt;/p&gt;

&lt;p&gt;Two details make this bad for container platforms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ordinary containers can reach it.&lt;/strong&gt; Creating Unix sockets and passing file descriptors are everyday operations. Docker's default seccomp profile allows them, and so does the &lt;code&gt;RuntimeDefault&lt;/code&gt; profile of the common container runtimes, because almost all software needs them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The kernel is shared.&lt;/strong&gt; A container is a process on the host kernel. Namespaces, cgroups and dropped capabilities limit what a process can do through normal interfaces, and a kernel memory bug like this one goes around them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Upstream fixed it in August (&lt;code&gt;af_unix: Unlink scc_entry in unix_del_edge()&lt;/code&gt;), and the kernel CVE team published the CVE on August 26. The fix is one line. Getting it into every distribution kernel is the slow part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who is affected
&lt;/h2&gt;

&lt;p&gt;According to the kernel CVE record, the vulnerable code is in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;every kernel from &lt;strong&gt;6.10&lt;/strong&gt; onward, until the fixed releases below&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;6.1.141 and later&lt;/strong&gt; in the 6.1 series, and &lt;strong&gt;6.6.93 and later&lt;/strong&gt; in the 6.6 series, because the new garbage collector was backported to those long-term branches&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix is in &lt;strong&gt;6.12.111&lt;/strong&gt; , &lt;strong&gt;6.18.53&lt;/strong&gt; , &lt;strong&gt;7.1.10&lt;/strong&gt; and &lt;strong&gt;7.2&lt;/strong&gt;. The CVE record lists no fixed 6.1 or 6.6 release, and a search of the 6.1.y and 6.6.y stable branch history on October 2 found no commit with the fix's title (the same search finds it in 6.12.y and 6.18.y). The newest releases there (6.1.188 and 6.6.157, both from September 14) still have the bug.&lt;/p&gt;

&lt;p&gt;What the main distributions say, checked on October 2:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Distribution&lt;/th&gt;
&lt;th&gt;Kernel&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu 26.04 LTS&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;linux&lt;/code&gt; and the cloud kernels (aws, azure, gcp, gke, oracle)&lt;/td&gt;
&lt;td&gt;Vulnerable, work in progress&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu 24.04 LTS&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;linux&lt;/code&gt; and the cloud kernels&lt;/td&gt;
&lt;td&gt;Vulnerable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu 22.04 LTS&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;linux&lt;/code&gt; (5.15)&lt;/td&gt;
&lt;td&gt;Not affected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debian 13 (trixie)&lt;/td&gt;
&lt;td&gt;6.12&lt;/td&gt;
&lt;td&gt;Fixed in &lt;code&gt;6.12.111-1&lt;/code&gt; (DSA-6528-1)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debian 12 (bookworm)&lt;/td&gt;
&lt;td&gt;6.1&lt;/td&gt;
&lt;td&gt;Vulnerable (&lt;code&gt;6.1.176-1&lt;/code&gt;), no fix yet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debian 11 (bullseye)&lt;/td&gt;
&lt;td&gt;5.10&lt;/td&gt;
&lt;td&gt;Not affected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Linux 2023&lt;/td&gt;
&lt;td&gt;6.18.35&lt;/td&gt;
&lt;td&gt;Livepatch for &lt;code&gt;kernel-livepatch-6.18.35-68.129&lt;/code&gt;, September 30 (ALAS2023LIVEPATCH-2026-397)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: the &lt;a href="https://ubuntu.com/security/CVE-2026-80521" rel="noopener noreferrer"&gt;Ubuntu CVE page&lt;/a&gt;, the &lt;a href="https://security-tracker.debian.org/tracker/CVE-2026-80521" rel="noopener noreferrer"&gt;Debian security tracker&lt;/a&gt; and the &lt;a href="https://alas.aws.amazon.com/AL2023/ALAS2023LIVEPATCH-2026-397.html" rel="noopener noreferrer"&gt;Amazon Linux advisory&lt;/a&gt;. We did not check every vendor. If you run RHEL and its rebuilds, Azure Linux, Container-Optimized OS, Bottlerocket or Talos, check that vendor's tracker.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some early reports listed Ubuntu 22.04 as vulnerable. Ubuntu's own tracker marks the 22.04 &lt;code&gt;linux&lt;/code&gt; kernel (5.15) as not affected, which matches the upstream record: 5.15 never received the new garbage collector. The 22.04 HWE kernels are newer and are tracked separately, so check them on the tracker.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why the version number is not enough
&lt;/h2&gt;

&lt;p&gt;The natural first step is to compare &lt;code&gt;uname -r&lt;/code&gt; against the list above. That works for upstream kernels and fails for distribution kernels, in both directions.&lt;/p&gt;

&lt;p&gt;Ubuntu 24.04 ships a 6.8 kernel. Upstream 6.8 never had the new garbage collector, so a version check says "not affected". Ubuntu's tracker says vulnerable: distribution kernels pull in fixes from the stable branches, and the backport that brought the new code into 6.6 can bring it into a distro's 6.8 too. The reverse also happens. A distro can backport the one-line fix and keep the old version number.&lt;/p&gt;

&lt;p&gt;So check for the code itself. The new garbage collector adds functions named &lt;code&gt;unix_add_edges&lt;/code&gt; and &lt;code&gt;unix_del_edges&lt;/code&gt;, and the kernel lists its function names in &lt;code&gt;/proc/kallsyms&lt;/code&gt;. That file is world-readable by default, and unprivileged users usually see the names with the addresses zeroed. An ordinary container sees the host kernel's list, because there is only one kernel (a gVisor or Kata sandbox does not).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-cE&lt;/span&gt; &lt;span class="s1"&gt;' unix_(add|del)_edges$'&lt;/span&gt; /proc/kallsyms

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;0&lt;/code&gt;: the new garbage collector was not found. On a standard distribution kernel, that means this CVE does not apply. On a custom or heavily modified kernel, confirm with whoever builds it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;1&lt;/code&gt; or more: the new garbage collector is present. Now the version and your vendor decide whether the fix is in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The presence check cannot see the fix itself, because the fix adds one line to a different function, &lt;code&gt;unix_del_edge()&lt;/code&gt;. That is why the script below combines it with the version, and sends you to your vendor when the kernel is a distro build. Treat it as a fast first filter, not a certificate.&lt;/p&gt;

&lt;h2&gt;
  
  
  A check script
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="c"&gt;# Heuristic check for CVE-2026-80521 (AF_UNIX garbage collector use-after-free).&lt;/span&gt;
&lt;span class="c"&gt;# Usually works inside a container too: /proc/kallsyms lists the host kernel's symbols.&lt;/span&gt;
&lt;span class="c"&gt;# Exit codes: 0 not detected or fixed, 1 affected, 2 check your vendor, 3 cannot tell.&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;

&lt;span class="nv"&gt;release&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;KERNEL_RELEASE&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="c"&gt;# KERNEL_RELEASE only for testing the logic&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[!&lt;/span&gt; &lt;span class="nv"&gt;$release&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;~ ^&lt;span class="o"&gt;([&lt;/span&gt;0-9]+&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="o"&gt;([&lt;/span&gt;0-9]+&lt;span class="o"&gt;)(&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="o"&gt;([&lt;/span&gt;0-9]+&lt;span class="o"&gt;))&lt;/span&gt;?&lt;span class="o"&gt;(&lt;/span&gt;.&lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: CANNOT TELL (unrecognised kernel release '&lt;/span&gt;&lt;span class="nv"&gt;$release&lt;/span&gt;&lt;span class="s2"&gt;')"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;3
&lt;span class="k"&gt;fi
&lt;/span&gt;&lt;span class="nv"&gt;major&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BASH_REMATCH&lt;/span&gt;&lt;span class="p"&gt;[1]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="nv"&gt;minor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BASH_REMATCH&lt;/span&gt;&lt;span class="p"&gt;[2]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="nv"&gt;patch&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BASH_REMATCH&lt;/span&gt;&lt;span class="p"&gt;[4]&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;0&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;
&lt;span class="nv"&gt;suffix&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BASH_REMATCH&lt;/span&gt;&lt;span class="p"&gt;[5]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="c"&gt;# anything after x.y.z: a distro, custom or -rc build&lt;/span&gt;

&lt;span class="c"&gt;# The bug is in the garbage collector that Linux 6.10 rewrote (backported to&lt;/span&gt;
&lt;span class="c"&gt;# 6.1.141+ and 6.6.93+). unix_add_edges and unix_del_edges exist only in that code.&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 1 /proc/kallsyms 2&amp;gt;/dev/null | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; .&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nv"&gt;new_gc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;unknown
&lt;span class="k"&gt;else
  &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-qE&lt;/span&gt; &lt;span class="s1"&gt;' unix_(add|del)_edges$'&lt;/span&gt; /proc/kallsyms
  &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="nv"&gt;$?&lt;/span&gt; &lt;span class="k"&gt;in &lt;/span&gt;0&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nv"&gt;new_gc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt; 1&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nv"&gt;new_gc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;no &lt;span class="p"&gt;;;&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nv"&gt;new_gc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;unknown &lt;span class="p"&gt;;;&lt;/span&gt; &lt;span class="k"&gt;esac&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Upstream releases that contain the fix (kernel CVE record, 2 October 2026).&lt;/span&gt;
&lt;span class="nv"&gt;fixed_upstream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;no
&lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$major&lt;/span&gt;&lt;span class="s2"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;$minor&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in
  &lt;/span&gt;6.12&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;((&lt;/span&gt; patch &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; 111 &lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nv"&gt;fixed_upstream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt;
  6.18&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;((&lt;/span&gt; patch &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; 53 &lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nv"&gt;fixed_upstream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt;
  7.1&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;((&lt;/span&gt; patch &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; 10 &lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nv"&gt;fixed_upstream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt;
&lt;span class="k"&gt;esac&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;((&lt;/span&gt; major &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; 7 &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;major &lt;span class="o"&gt;==&lt;/span&gt; 7 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; minor &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; 2&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;))&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then &lt;/span&gt;&lt;span class="nv"&gt;fixed_upstream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"kernel: &lt;/span&gt;&lt;span class="nv"&gt;$release&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"new AF_UNIX GC found: &lt;/span&gt;&lt;span class="nv"&gt;$new_gc&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="nv"&gt;$new_gc&lt;/span&gt; &lt;span class="k"&gt;in
  &lt;/span&gt;unknown&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: CANNOT TELL (could not read /proc/kallsyms); check &lt;/span&gt;&lt;span class="nv"&gt;$major&lt;/span&gt;&lt;span class="s2"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;$minor&lt;/span&gt;&lt;span class="s2"&gt; with your vendor"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;3 &lt;span class="p"&gt;;;&lt;/span&gt;
  no&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: NOT DETECTED (no new garbage collector; on a standard distro kernel this means not affected)"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;0 &lt;span class="p"&gt;;;&lt;/span&gt;
&lt;span class="k"&gt;esac&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nv"&gt;$suffix&lt;/span&gt;&lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt;&lt;span class="nv"&gt;$fixed_upstream&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nv"&gt;$suffix&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="nt"&gt;-rc&lt;/span&gt;&lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: PROBABLY FIXED (base &lt;/span&gt;&lt;span class="nv"&gt;$major&lt;/span&gt;&lt;span class="s2"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;$minor&lt;/span&gt;&lt;span class="s2"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;$patch&lt;/span&gt;&lt;span class="s2"&gt; has the fix upstream); confirm with your vendor"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;2
  &lt;span class="k"&gt;fi
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: CHECK YOUR VENDOR (new garbage collector present; distro kernels backport fixes without changing the version)"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;2
&lt;span class="k"&gt;fi
if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt;&lt;span class="nv"&gt;$fixed_upstream&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nb"&gt;yes&lt;/span&gt;&lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: FIXED (if this is an unmodified upstream &lt;/span&gt;&lt;span class="nv"&gt;$release&lt;/span&gt;&lt;span class="s2"&gt; build)"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="k"&gt;fi
&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"verdict: AFFECTED (upstream &lt;/span&gt;&lt;span class="nv"&gt;$release&lt;/span&gt;&lt;span class="s2"&gt; has the bug and no fix)"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We ran it as an unprivileged user on two machines: a Raspberry Pi test box on Raspberry Pi OS, and an Ubuntu 22.04 server. The &lt;code&gt;KERNEL_RELEASE&lt;/code&gt; variable is only there to test the version logic with made-up strings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;af-unix-gc-check&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Raspberry Pi OS, kernel 6.12.62: new garbage collector present, below the 6.12.111 fix
$ ./af-unix-gc-check.sh; echo "exit=$?"
kernel: 6.12.62+rpt-rpi-v8
new AF_UNIX GC found: yes
verdict: CHECK YOUR VENDOR (new garbage collector present; distro kernels backport fixes without changing the version)
exit=2
# Ubuntu 22.04, kernel 5.15: old garbage collector
$ ./af-unix-gc-check.sh; echo "exit=$?"
kernel: 5.15.0-191-generic
new AF_UNIX GC found: no
verdict: NOT DETECTED (no new garbage collector; on a standard distro kernel this means not affected)
exit=0

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We also read &lt;code&gt;/proc/kallsyms&lt;/code&gt; from inside a running Docker container on the 22.04 server. The symbol names were there (addresses zeroed), so the one-line check worked from inside the container. Hardened setups can mask that file; then the script says it cannot tell.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check a whole cluster
&lt;/h2&gt;

&lt;p&gt;Start with the kernel and OS image of every node:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get nodes &lt;span class="nt"&gt;-o&lt;/span&gt; custom-columns&lt;span class="o"&gt;=&lt;/span&gt;NAME:.metadata.name,OS:.status.nodeInfo.osImage,KERNEL:.status.nodeInfo.kernelVersion

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then run the presence check on every distinct kernel build you see, or simply on every node during a rollout, when old and new images run side by side:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Starts a debugging pod on the node; /proc/kallsyms is the node's kernel symbol list&lt;/span&gt;
kubectl debug node/&amp;lt;node-name&amp;gt; &lt;span class="nt"&gt;-it&lt;/span&gt; &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;busybox:1.36 &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-cE&lt;/span&gt; &lt;span class="s1"&gt;' unix_(add|del)_edges$'&lt;/span&gt; /proc/kallsyms

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;kubectl debug node&lt;/code&gt; creates a pod that shares the node's PID, network and IPC namespaces, so it needs the matching RBAC and must be allowed by your admission policy. Delete the debug pod when you are done (&lt;code&gt;kubectl get pods&lt;/code&gt; shows it as &lt;code&gt;node-debugger-...&lt;/code&gt;).&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Node kernel&lt;/strong&gt; uname -r&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;kallsyms check&lt;/strong&gt; unix_add_edges?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix in this build?&lt;/strong&gt; version or vendor tracker&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Outcomes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;0 matches or fixed build: not exposed&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code present, no fix: isolate untrusted workloads&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What to do now
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Patch where a fix exists
&lt;/h3&gt;

&lt;p&gt;If your kernel line has a fix (Debian 13, upstream 6.12.111+, 6.18.53+, 7.1.10+ or 7.2+), install it. A new kernel package does nothing until the node boots into it, so plan the reboots: cordon and drain one node at a time, or replace node pools with a new image. On Amazon Linux 2023 with the matching kernel, the livepatch applies without a reboot once kernel livepatching is enabled; verify that it is applied.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Find the workloads that run code you did not write
&lt;/h3&gt;

&lt;p&gt;Where there is no fix yet (Ubuntu 24.04 and 26.04, Debian 12, anything on a 6.1 or 6.6 kernel at or above the backport), the question is who can run arbitrary code on those nodes. A container escape needs code execution inside a container first. The usual places where strangers get exactly that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CI runners and build agents that run pull or merge request code from forks&lt;/li&gt;
&lt;li&gt;multi-tenant clusters where teams or customers deploy their own images&lt;/li&gt;
&lt;li&gt;notebooks, online code runners, plugin systems and "run your script" features&lt;/li&gt;
&lt;li&gt;any pod with a known remote code execution bug that you have not patched yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Single-tenant nodes that only run your own images are still worth patching, but they are not where you start.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Move untrusted workloads into a sandbox or onto unaffected nodes
&lt;/h3&gt;

&lt;p&gt;For the workloads in step 2, the most reliable option you can use today is a different isolation boundary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;gVisor&lt;/strong&gt; runs the container on a user-space kernel (the Sentry) that implements Unix sockets itself, so the sandbox's own Unix sockets do not use the host kernel's AF_UNIX garbage collector. (Access to host Unix sockets is a separate option and is off by default.) On GKE this is GKE Sandbox. Elsewhere, install &lt;code&gt;runsc&lt;/code&gt; on the nodes, register it as a containerd runtime handler, and then add a RuntimeClass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kata Containers&lt;/strong&gt; runs each pod in a lightweight VM, and can use Firecracker as its VM monitor. The guest kernel may have the same bug, but an escape then lands in a VM, not on the host.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A separate node pool on an unaffected kernel&lt;/strong&gt; , such as Ubuntu 22.04 with its 5.15 kernel, for the untrusted workloads until your main image is fixed. It is a short-term measure with its own tradeoffs (an older kernel and different hardware support), not a destination.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A RuntimeClass for gVisor, and a pod that uses it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;node.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RuntimeClass&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gvisor&lt;/span&gt;
&lt;span class="na"&gt;handler&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;runsc&lt;/span&gt; &lt;span class="c1"&gt;# the containerd runtime handler configured on the node&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;untrusted-build&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;runtimeClassName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gvisor&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;build&lt;/span&gt;
      &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;registry.example.com/ci/build-runner:2026.10&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Test your workloads under gVisor before you switch: it does not implement every system call, and I/O heavy jobs can run slower.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Know what does not help
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default seccomp.&lt;/strong&gt; The default profiles allow Unix sockets and file descriptor passing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blocking AF_UNIX with a custom seccomp profile, in general.&lt;/strong&gt; For the &lt;a href="https://devops-daily.com/posts/copy-fail-cve-2026-31431-linux-container-escape" rel="noopener noreferrer"&gt;Copy Fail AF_ALG bug&lt;/a&gt; earlier this year, blocking one rarely used socket family was a clean mitigation. AF_UNIX is different: DNS resolution through NSS, database clients on local sockets, process managers, logging and many language runtimes use it. Blocking it breaks most real applications. A custom profile can still make sense for a specific workload that you have tested without Unix sockets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rootless containers, user namespaces and dropped capabilities.&lt;/strong&gt; They are good hardening, but they are not a fix here: the vulnerable code is reachable without any special privilege.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Watch for the fix and re-check
&lt;/h3&gt;

&lt;p&gt;Subscribe to your distribution's security notices, and keep the check above in your node image pipeline. Once the fixed kernel is out, the version check plus the vendor's fixed package version tells you which nodes still need a reboot. Then move the sandboxed workloads back, or keep them sandboxed. Copy Fail and Fragnesia were container escapes too, earlier this year, so the sandbox may be worth keeping.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;CVE-2026-80521 is a use-after-free in the Linux AF_UNIX garbage collector. A public exploit escapes default Docker and Kubernetes containers to host root.&lt;/li&gt;
&lt;li&gt;It is not limited to Ubuntu. The vulnerable code is in 6.10 and later and in the 6.1.141+ and 6.6.93+ long-term kernels. Upstream fixed 6.12, 6.18 and 7.1, but not 6.1 or 6.6.&lt;/li&gt;
&lt;li&gt;Ubuntu 24.04 and 26.04 and Debian 12 had no fixed kernel on October 2. Ubuntu 22.04 (5.15), Debian 11 and other kernels without the new garbage collector are not affected.&lt;/li&gt;
&lt;li&gt;Version numbers mislead on distro kernels. &lt;code&gt;grep -cE ' unix_(add|del)_edges$' /proc/kallsyms&lt;/code&gt; tells you whether the new garbage collector is there, usually even from inside a container.&lt;/li&gt;
&lt;li&gt;Until you can patch, put workloads that run untrusted code behind gVisor, Kata or a VM, or on nodes with an unaffected kernel. Default seccomp and user namespaces do not stop this one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://devops-daily.com/posts/af-unix-container-escape-cve-2026-80521" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>linux</category>
      <category>security</category>
      <category>kubernetes</category>
      <category>containers</category>
    </item>
    <item>
      <title>Hacktoberfest 2026 Stopped Counting PRs. Make Your First Ones Anyway, One a Day</title>
      <dc:creator>DevOps Daily</dc:creator>
      <pubDate>Thu, 01 Oct 2026 15:04:37 +0000</pubDate>
      <link>https://dev.to/devopsdaily/hacktoberfest-2026-stopped-counting-prs-make-your-first-ones-anyway-one-a-day-377b</link>
      <guid>https://dev.to/devopsdaily/hacktoberfest-2026-stopped-counting-prs-make-your-first-ones-anyway-one-a-day-377b</guid>
      <description>&lt;p&gt;Hacktoberfest started today, and if you have taken part before, the first thing to know is that the rule everyone remembers is gone. Pull requests no longer count toward Hacktoberfest rewards. There is no PR target and no PR-based swag this year.&lt;/p&gt;

&lt;p&gt;That is a good change. It is also a slightly awkward one if Hacktoberfest was the push you needed to make your first open source contribution. This post covers what changed, why a first PR is still worth making, and a 7-day challenge we built: one small task a day, 5 to 15 minutes each, and each one ends with a real pull request that a maintainer reviews.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in Hacktoberfest 2026
&lt;/h2&gt;

&lt;p&gt;From the &lt;a href="https://hacktoberfest.com/questions/" rel="noopener noreferrer"&gt;official FAQ&lt;/a&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Pull requests and merge requests will no longer count toward Hacktoberfest rewards. It’s easier than ever to submit low-effort spam PRs to projects, so we’re listening to maintainer feedback and no longer actively incentivizing PRs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The rest of the event changed with it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;New organizers.&lt;/strong&gt; MLH and DEV run Hacktoberfest this year, in partnership with DigitalOcean.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New format.&lt;/strong&gt; 300+ in-person and online community events, focused on hands-on building and learning with open-source AI and open-weight models. The in-person ones are called Fests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New rewards.&lt;/strong&gt; You collect stickers. Two are required: sign in with MyMLH and add your mailing address. You earn more by checking in at a Fest, entering the code shown during a livestream, or submitting to a DEV Challenge. Fifteen stickers enter you in a raffle for a Hacktoberfest t-shirt or an Arduino Uno Q board (details on the &lt;a href="https://hacktoberfest.com/online/" rel="noopener noreferrer"&gt;online participation page&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you maintain a repo, you know why. Every October, maintainers spent hours closing pull requests that changed one word in a README, and AI tools made those PRs even cheaper to produce. Taking the swag off the PR removes the main reason to send them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why your first PR still matters
&lt;/h2&gt;

&lt;p&gt;The spam was the problem, not the pull request. The things a first contribution teaches you did not change:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Forking, branching and keeping a fork in sync without breaking it.&lt;/li&gt;
&lt;li&gt;Reading a codebase you did not write, well enough to change one thing in it.&lt;/li&gt;
&lt;li&gt;Running a project's tests before a reviewer has to tell you they fail.&lt;/li&gt;
&lt;li&gt;Writing a PR description that a busy maintainer can approve in one read.&lt;/li&gt;
&lt;li&gt;Taking review feedback and pushing a fix to the same branch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are daily skills in any DevOps or platform job. A merged PR in a public repo is also evidence of them, which a line on a CV is not.&lt;/p&gt;

&lt;p&gt;What you need is a project that wants small contributions, says exactly how to make them, and reviews them. So we made one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The challenge: one small PR a day for 7 days
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://devops-daily.com/hacktoberfest" rel="noopener noreferrer"&gt;DevOps Daily Hacktoberfest challenge&lt;/a&gt; is 7 daily tasks, plus a bonus, in the &lt;a href="https://github.com/The-DevOps-Daily/devops-daily" rel="noopener noreferrer"&gt;devops-daily repo&lt;/a&gt;: an open source DevOps learning site with 560+ posts, 45 quizzes and 27 flashcard decks, all stored as Markdown and JSON in the repo. Each day has its own page with the file to change and an example, and most days include a command to check your work.&lt;/p&gt;

&lt;p&gt;To be clear about what this is: it is a DevOps Daily challenge that runs during Hacktoberfest, not an official Hacktoberfest event, and these PRs do not count toward Hacktoberfest stickers. What you get is a real contribution, reviewed by a maintainer, with a test suite in CI, on your GitHub profile. The first challenge PRs came in within hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set up once (about 10 minutes)
&lt;/h2&gt;

&lt;p&gt;You need Git, Node.js 22.13.1 or later (below 25), and pnpm 10.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Fork https://github.com/The-DevOps-Daily/devops-daily on GitHub first, then:&lt;/span&gt;
git clone https://github.com/&amp;lt;your-username&amp;gt;/devops-daily.git
&lt;span class="nb"&gt;cd &lt;/span&gt;devops-daily
git remote add upstream https://github.com/The-DevOps-Daily/devops-daily.git

pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;pnpm &lt;span class="nb"&gt;test&lt;/span&gt;          &lt;span class="c"&gt;# the unit tests CI runs on every PR&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before each day's task, start a new branch from a fresh copy of main, so your PR contains only that day's change. Change the day number each time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git fetch upstream
git checkout &lt;span class="nt"&gt;-b&lt;/span&gt; hacktoberfest/day-1 upstream/main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The one exception is Day 7, which adds to your Day 1 profile. If your Day 1 PR is not merged yet, branch from your Day 1 branch instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 7 days
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Day 1: Add yourself to the experts directory (5 min).&lt;/strong&gt; Create a Markdown profile in &lt;code&gt;content/experts/&lt;/code&gt; with your name, a short bio and how people can reach you. It is the smallest possible PR, and it shows you the whole fork, branch, commit, push and PR flow once before the content gets harder.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 2: Add your favorite DevOps tool (5 min).&lt;/strong&gt; Add one entry to the toolbox page: name, one-line description, link and category. Pick a tool you use, not one you have heard of. The description is where your experience shows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 3: Add a quiz question (5 min).&lt;/strong&gt; Add a question to one of the existing quizzes in &lt;code&gt;content/quizzes/&lt;/code&gt;. The tests check every question's shape, so run them before you push. Write the explanation for the wrong answers too: "why not B" teaches more than "why A".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 4: Add a flashcard (5 min).&lt;/strong&gt; Add one or two cards to a deck in &lt;code&gt;content/flashcards/&lt;/code&gt;. A good card has one fact on the front and the shortest complete answer on the back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 5: Share a tip (10 min).&lt;/strong&gt; Add a gotcha or "thing I wish I knew" to an existing post or guide. This is the first day where you edit someone else's writing, so add your tip where a reader needs it and leave the rest of the post alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 6: Find and fix something (10 min).&lt;/strong&gt; Browse &lt;a href="https://devops-daily.com" rel="noopener noreferrer"&gt;devops-daily.com&lt;/a&gt; and find a typo, a broken link, a deprecated command or a formatting problem. Fix exactly one thing, and put the URL where you found it in the PR description.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 7: Share your stack (15 min).&lt;/strong&gt; Add a "My Stack" section to the profile you made on Day 1: what you run, why you picked it, and what you would change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bonus: Build something (30+ min).&lt;/strong&gt; A full quiz on a topic that is not covered yet, a tool comparison, a production checklist, or for the ambitious, an interactive game or simulator.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to get your PR merged (and not closed as spam)
&lt;/h2&gt;

&lt;p&gt;The 2026 change is a good reminder of what maintainers are tired of. These habits make a PR easy to review in any project, not only this one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One change per PR.&lt;/strong&gt; If you spot a typo in another quiz while doing Day 3, fix it in a separate PR. A small diff is quick to check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run the tests first.&lt;/strong&gt; Run &lt;code&gt;pnpm test&lt;/code&gt; before you push, so you find a failing check before a reviewer does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Say what and why.&lt;/strong&gt; Two sentences: what you changed, and why it is correct. For a fix, add the link to where you found the problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read what you submit.&lt;/strong&gt; If an AI tool helped you write a quiz question or a tip, check it like you would check a stranger's code. A confident wrong answer in a quiz is worse than no question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer review on the same branch.&lt;/strong&gt; Push the fix to the same branch and the PR updates. Do not open a new PR for each round.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Collect your Hacktoberfest stickers too
&lt;/h2&gt;

&lt;p&gt;The challenge and the official event go together well. Do one small PR a day here, and collect your stickers at &lt;a href="https://hacktoberfest.com/" rel="noopener noreferrer"&gt;hacktoberfest.com&lt;/a&gt;: check in at a Fest near you, enter the code during a livestream, or submit to a DEV Challenge, since DEV co-runs the event this year.&lt;/p&gt;

&lt;p&gt;Start with Day 1 today: &lt;a href="https://devops-daily.com/hacktoberfest" rel="noopener noreferrer"&gt;devops-daily.com/hacktoberfest&lt;/a&gt;. And if you get stuck on any day, open an issue on the repo or ask in your PR. Questions are contributions too.&lt;/p&gt;

</description>
      <category>hacktoberfest</category>
      <category>opensource</category>
      <category>beginners</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
