<?xml version="1.0" encoding="utf-8" standalone="yes"?><?xml-stylesheet href="/feed.css?v=67325d779b74" type="text/css"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:site="https://lalitm.com/feed/ns#"><channel><title>Lalit Maganti (Tag: Devtools)</title><link>https://lalitm.com/tags/devtools/</link><description>Recent content tagged Devtools on Lalit Maganti</description><site:notice>This is a feed.
Feeds let you subscribe to updates from this site using a feed reader. Copy this page's URL from your address bar and paste it into your reader.
New to feeds? Read: https://aboutfeeds.com</site:notice><docs>https://aboutfeeds.com</docs><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Sat, 12 Sep 2026 14:38:00 +0100</lastBuildDate><atom:link href="https://lalitm.com/tags/devtools/index.xml" rel="self" type="application/rss+xml"/><item><title>I made a build visualizer to understand Bun’s compile times</title><link>https://lalitm.com/post/buildprof/</link><pubDate>Sat, 12 Sep 2026 14:38:00 +0100</pubDate><guid>https://lalitm.com/post/buildprof/</guid><description>I built buildprof (Github), an open-source tracing tool that shows where the time goes when you compile software on Linux. Here’s a realtime video of it profiling a clean build of ripgrep:
Watch the buildprof demo
Sometimes, builds are slow because there is simply a lot of code to compile. But more often than not, there are fixable problems: poor parallelism, repeated work, dependency downloads or a huge compiler/linker invocation. buildprof makes all of this clearly visible, so you can see what’s worth investigating and optimizing.
You run it by putting buildprof -- in front of any build command you already use:
buildprof -- make -j16 buildprof -- cargo build buildprof -- ninja -C out/target buildprof -- just build buildprof -- ./dev/custom-build-script.sh buildprof records every process your build command launches, including their subprocesses (and their subprocesses…), and lays them out on one timeline. Time moves from left to right, bar width shows duration, and child processes appear beneath whatever launched them.</description><content:encoded>&lt;p&gt;I built &lt;a href="https://buildprof.lalitm.com"&gt;buildprof&lt;/a&gt;
(&lt;a href="https://github.com/lalitMaganti/buildprof"&gt;Github&lt;/a&gt;), an open-source tracing
tool that shows where the time goes when you compile software on Linux. Here&amp;rsquo;s a
realtime video of it profiling a clean build of ripgrep:&lt;/p&gt;
&lt;p&gt;&lt;video src="https://lalitm.com/assets/buildprof/buildprof-initial-demo-v2.mp4" autoplay muted loop playsinline style="width: 100%;"&gt;&lt;a href="https://lalitm.com/assets/buildprof/buildprof-initial-demo-v2.mp4"&gt;Watch the buildprof demo&lt;/a&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Sometimes, builds are slow because there is simply a lot of code to compile. But
more often than not, there are fixable problems: poor parallelism, repeated
work, dependency downloads or a huge compiler/linker invocation. buildprof makes
all of this clearly visible, so you can see what’s worth investigating and
optimizing.&lt;/p&gt;
&lt;p&gt;You run it by putting &lt;code&gt;buildprof --&lt;/code&gt; in front of any build command you already
use:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;buildprof -- make -j16
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;buildprof -- cargo build
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;buildprof -- ninja -C out/target
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;buildprof -- just build
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;buildprof -- ./dev/custom-build-script.sh
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;buildprof records every process your build command launches, including their
subprocesses (and their subprocesses&amp;hellip;), and lays them out on one timeline.
Time moves from left to right, bar width shows duration, and child processes
appear beneath whatever launched them.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/ripgrep-overview.png" alt="A ripgrep build: Cargo spans the whole build, rustc invocations compile crates in parallel, and the final rustc invocation launches a linker chain."&gt;&lt;/p&gt;
&lt;p&gt;I made buildprof because
&lt;a href="https://x.com/jarredsumner/status/2090619419059974620"&gt;this tweet&lt;/a&gt; from Jarred
Sumner, chief architect of the Bun JavaScript runtime, was living rent free in
my head:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/jarred-bun-rust-compile-times.png" alt="Jarred Sumner&amp;rsquo;s comparison showing a 30 minute 6 second median Linux build for Bun 1.3.14 and 5 minute 37 second median for Bun 1.4.0."&gt;&lt;/p&gt;
&lt;p&gt;Specifically, the claim that Bun’s new Rust build was &amp;gt;5× faster on Linux than
its old Zig build really bothered me. In my experience, Zig projects had usually
compiled &lt;em&gt;much&lt;/em&gt; faster than Rust projects of similar complexity. That intuition
was enough to make me feel there was a mystery to solve.&lt;/p&gt;
&lt;p&gt;This was further compounded by another important, yet easily missed, detail in
the tweet: the Zig build used Full LTO, while the Rust build used ThinLTO.&lt;/p&gt;
&lt;p&gt;Compilers normally optimize separate compilation units largely in
isolation.&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt; Link-time optimization (LTO) lets them optimize
&lt;strong&gt;across&lt;/strong&gt; those boundaries. Full LTO brings those units together into one large
optimization job, while ThinLTO preserves more separation so much of the work
can run in parallel.&lt;/p&gt;
&lt;p&gt;From past experience, this difference can have an &lt;strong&gt;enormous&lt;/strong&gt; effect on build
time. The tweet mentioned it in passing, but I wondered how much of the headline
improvement it explained.&lt;/p&gt;
&lt;p&gt;I started by trying to reproduce the numbers.&lt;/p&gt;
&lt;h2 id="the-numbers-reproduced-but-now-what"&gt;The numbers reproduced. But now what?&lt;a class="heading-anchor" href="#the-numbers-reproduced-but-now-what" aria-label="Permalink to The numbers reproduced. But now what?"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;I checked out
&lt;a href="https://github.com/oven-sh/bun/releases/tag/bun-v1.3.14"&gt;Bun 1.3.14&lt;/a&gt; and
&lt;a href="https://github.com/oven-sh/bun/releases/tag/bun-v1.4.0"&gt;Bun 1.4.0&lt;/a&gt; and wrote
&lt;a href="https://github.com/LalitMaganti/blog-code/tree/main/buildprof-bun"&gt;some scripts&lt;/a&gt;
to replay their Linux x64 CI builds on a 6-core, 12-thread Linux VM. The scripts
preserved the build steps and their dependencies, running everything on one
machine.&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;My timings were in the same ballpark as Jarred’s:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Linux x64 build&lt;/th&gt;
&lt;th style="text-align: right"&gt;Zig era&lt;/th&gt;
&lt;th style="text-align: right"&gt;Rust era&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bun&amp;rsquo;s reported CI median&lt;/td&gt;
&lt;td style="text-align: right"&gt;30m06s&lt;/td&gt;
&lt;td style="text-align: right"&gt;5m37s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;My single-machine CI-profile replay&lt;/td&gt;
&lt;td style="text-align: right"&gt;24m24s&lt;/td&gt;
&lt;td style="text-align: right"&gt;5m40s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;OK, so the gap showed up on my machine too. But a lot had changed between the
two measurements besides the language; so what was actually responsible? Was it
the Zig compiler that was taking all that extra time? Or maybe it was the Full
LTO link? Or perhaps there was something else in Bun’s build I hadn’t even
thought to look at?&lt;/p&gt;
&lt;p&gt;This is where my profiling and developer-tools brain kicked in. Usually, when
I’m trying to understand why something is slow, I want a trace: what happened,
when it happened and how long it took. It would be really cool to have that for
these builds, to put them on a timeline and see where their time actually went.&lt;/p&gt;
&lt;p&gt;But a build involves a lot of different tools, each with its own idea of what’s
happening. What could I record that would let me see across all of them?&lt;/p&gt;
&lt;h2 id="builds-are-process-trees"&gt;Builds are process trees&lt;a class="heading-anchor" href="#builds-are-process-trees" aria-label="Permalink to Builds are process trees"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;When you type &lt;code&gt;cargo build&lt;/code&gt; or &lt;code&gt;zig build&lt;/code&gt;, it feels like you are running one
program. The build system works out what needs to be rebuilt, the ordering
between those pieces and what can run in parallel. But generally, it does not
perform all that work itself; it launches compilers, code generators, archivers,
linkers and arbitrary scripts. Which can launch more programs which launch some
more…&lt;/p&gt;
&lt;p&gt;Different build systems describe that work in different ways. Cargo sees crates,
Ninja sees build edges and CMake generates instructions for another build
system. From the operating system&amp;rsquo;s point of view, however, they (mostly) look
like processes launching other processes.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;A Rust build, for example, might contain a chain like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;cargo
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;└── rustc
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; └── cc
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; └── collect2
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; └── ld.lld
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If we record when each subprocess starts and ends, we can lay them out on a
timeline. Here’s what that chain looks like in buildprof:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/ripgrep-process-chain.png" alt="The final link in a ripgrep build, showing cargo launching rustc, then cc, collect2 and ld.lld beneath it."&gt;&lt;/p&gt;
&lt;p&gt;There are also several nice properties to visualizing a build at this layer:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;It&amp;rsquo;s build-system agnostic&lt;/strong&gt;: Cargo, Ninja, Zig, Make and most other build
systems do much of their work by spawning processes, so we do not need to
write a special integration for each one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It naturally includes custom scripts&lt;/strong&gt;: This includes both scripts above
the build system (repository setup, dependency fetching) and scripts
underneath it (code generators, asset processors).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;We can follow the files between build steps&lt;/strong&gt;: recording which files each
process reads and writes lets us see which steps produce the inputs for
others. This even works across build systems!&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This gave me a starting point for buildprof: record the process tree, then turn
it into a timeline I could explore. There are plenty more details to get into,
which I will do later. But once I had that working, I could finally go back to
my initial question: what was Bun doing for those twenty-four minutes?&lt;/p&gt;
&lt;h2 id="pointing-it-at-bun"&gt;Pointing it at Bun&lt;a class="heading-anchor" href="#pointing-it-at-bun" aria-label="Permalink to Pointing it at Bun"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="why-was-the-zig-ci-build-so-much-slower"&gt;Why was the Zig CI build so much slower?&lt;a class="heading-anchor" href="#why-was-the-zig-ci-build-so-much-slower" aria-label="Permalink to Why was the Zig CI build so much slower?"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;I started by recording the Zig-era CI build with buildprof, using the
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/scripts/record-original-zig-ci.sh"&gt;same scripts as before&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-ci-overview-notes.png" alt="The complete Zig-era CI build"&gt;&lt;/p&gt;
&lt;p style="text-align: center;"&gt;&lt;em&gt;&lt;a href="https://buildprof.lalitm.com/v0.2.2/#!/?url=https%3A%2F%2Fblogexamples.lalitm.com%2Fbuildprof-bun%2F2026-09-04%2Fbun-zig-original-ci-b5a45845003ecae0.buildprof"&gt;Explore in buildprof&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Right away we can see a huge problem: the &lt;code&gt;ld.lld&lt;/code&gt; linker invocation &lt;em&gt;dominates&lt;/em&gt;
the build time. It ran alone at the very end for over sixteen minutes, about
two-thirds of the entire build. What the heck was it doing for all that time?&lt;/p&gt;
&lt;p&gt;Clicking on the linker shows its command line, which buildprof captures
automatically:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-full-lto-command-notes.png" alt="The selected Zig-era linker and its Full LTO flag"&gt;&lt;/p&gt;
&lt;p&gt;There’s Full LTO, just as Jarred said. Given how long the link was taking, it
was now my main suspect.&lt;/p&gt;
&lt;p&gt;But the process tree alone couldn’t tell me whether LTO was actually responsible
for those sixteen minutes. Thankfully, LLD records its own internal timing
events, and buildprof can include them when you use &lt;code&gt;--compiler-traces&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;I
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/scripts/record-zig-full-lto-link-detail.sh"&gt;recorded the final link again&lt;/a&gt;,
this time with &lt;code&gt;--compiler-traces&lt;/code&gt; enabled:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-full-lto-linker-internals-notes.png" alt="LLD&amp;rsquo;s internal phases"&gt;&lt;/p&gt;
&lt;p style="text-align: center;"&gt;&lt;em&gt;&lt;a href="https://buildprof.lalitm.com/v0.2.2/#!/?url=https%3A%2F%2Fblogexamples.lalitm.com%2Fbuildprof-bun%2F2026-09-04%2Fbun-zig-full-lto-link-detail-d1e7a0e1dd7750d4.buildprof"&gt;Explore in buildprof&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Now we can see that LTO &lt;em&gt;is&lt;/em&gt; where almost all the time goes. The linker is
running compiler passes over the program, not just combining already-compiled
files. The &lt;code&gt;OptModule&lt;/code&gt; bar alone takes just over ten minutes and includes the
passes which generate machine code.&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;h4 id="how-did-the-rust-ci-build-differ"&gt;How did the Rust CI build differ?&lt;a class="heading-anchor" href="#how-did-the-rust-ci-build-differ" aria-label="Permalink to How did the Rust CI build differ?"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;With so much of the Zig build spent in LTO, I wanted to see how much time the
Rust build spent linking. I recorded that build too:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/rust-ci-overview-notes.png" alt="The complete Rust-era CI build"&gt;&lt;/p&gt;
&lt;p style="text-align: center;"&gt;&lt;em&gt;&lt;a href="https://buildprof.lalitm.com/v0.2.2/#!/?url=https%3A%2F%2Fblogexamples.lalitm.com%2Fbuildprof-bun%2F2026-09-04%2Fbun-rust-original-ci-4c9dbcbe6d041d98.buildprof"&gt;Explore in buildprof&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Just 2m24s. And this time, as expected, the linker command contains
&lt;code&gt;-plugin-opt=thinlto&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/rust-ci-thinlto-notes.png" alt="The Rust linker invocation with ThinLTO enabled"&gt;&lt;/p&gt;
&lt;p&gt;Both builds were doing LTO, but with different settings and very different link
times. What if I kept Bun’s Zig code and changed Full LTO to ThinLTO? How much
of the gap would that close?&lt;/p&gt;
&lt;h4 id="trying-thinlto"&gt;Trying ThinLTO&lt;a class="heading-anchor" href="#trying-thinlto" aria-label="Permalink to Trying ThinLTO"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;I
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/patches/zig-thinlto.patch"&gt;switched Zig Bun’s build flags to ThinLTO&lt;/a&gt;
and recorded another clean build, along with a fresh Full-LTO build for
comparison:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-lto-pair-comparison-notes.png" alt="The matched Full-LTO build and partial ThinLTO experiment"&gt;&lt;/p&gt;
&lt;p style="text-align: center;"&gt;&lt;em&gt;Explore in buildprof: &lt;a href="https://buildprof.lalitm.com/v0.2.2/#!/?url=https%3A%2F%2Fblogexamples.lalitm.com%2Fbuildprof-bun%2F2026-09-04%2Fbun-zig-ci-dag-full-lto-6111eb9fd9486a35.buildprof"&gt;Full LTO&lt;/a&gt; · &lt;a href="https://buildprof.lalitm.com/v0.2.2/#!/?url=https%3A%2F%2Fblogexamples.lalitm.com%2Fbuildprof-bun%2F2026-09-04%2Fbun-zig-ci-dag-thin-lto-fdfb05e3d506d810.buildprof"&gt;partial ThinLTO&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The link got 3m40s faster in this pair of recordings, but it was still taking
nearly thirteen minutes. Why was linking still so expensive?&lt;/p&gt;
&lt;p&gt;Looking back at the compiler trace, a lot of the work was on functions with
&lt;code&gt;JSC&lt;/code&gt; in their names. That’s JavaScriptCore, the engine Bun uses to execute
JavaScript. The linker was spending time compiling the JavaScript engine
too.&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Clicking on the linker invocation showed the
&lt;a href="https://github.com/oven-sh/WebKit/tree/5488984d20e0dbfe4be2c3ba8fb18eb81a5e0e8b"&gt;WebKit&lt;/a&gt;
libraries among its inputs, including &lt;code&gt;libJavaScriptCore.a&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-thinlto-webkit-inputs-notes.png" alt="The linker command has ThinLTO enabled but still includes WebKit’s libraries, including libJavaScriptCore.a."&gt;&lt;/p&gt;
&lt;p&gt;Following those inputs back through the build, I found that Bun wasn’t compiling
these libraries itself. It was downloading them from a separate WebKit build.
And when I checked
&lt;a href="https://github.com/oven-sh/WebKit/blob/5488984d20e0dbfe4be2c3ba8fb18eb81a5e0e8b/Dockerfile#L4"&gt;that build’s flags&lt;/a&gt;,
there it was again: &lt;code&gt;-flto=full&lt;/code&gt;. The Rust build used a newer WebKit revision
whose
&lt;a href="https://github.com/oven-sh/WebKit/blob/0f966e81b78c84bb/Dockerfile#L4"&gt;build recipe selected ThinLTO&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Even though I had changed how Bun compiled its own code, those downloaded
libraries still contained Full-LTO inputs and so the linker still had to
optimize that code and turn it into machine code. To change that, I would have
to rebuild WebKit too.&lt;/p&gt;
&lt;h4 id="rebuilding-webkit"&gt;Rebuilding WebKit&lt;a class="heading-anchor" href="#rebuilding-webkit" aria-label="Permalink to Rebuilding WebKit"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;I checked out the historical WebKit revision and rebuilt it and its ICU
dependencies with compatible ThinLTO settings. Then I replaced the downloaded
libraries with the ones I had built, keeping the ThinLTO changes to Bun.&lt;/p&gt;
&lt;p&gt;Here are the recorded builds:&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Zig-era build&lt;/th&gt;
&lt;th style="text-align: right"&gt;Whole build&lt;/th&gt;
&lt;th style="text-align: right"&gt;Final linker&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Original Full LTO&lt;/td&gt;
&lt;td style="text-align: right"&gt;24m24s&lt;/td&gt;
&lt;td style="text-align: right"&gt;16m35s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bun ThinLTO; original WebKit archives&lt;/td&gt;
&lt;td style="text-align: right"&gt;20m20s&lt;/td&gt;
&lt;td style="text-align: right"&gt;12m55s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bun ThinLTO; rebuilt ThinLTO WebKit and ICU&lt;/td&gt;
&lt;td style="text-align: right"&gt;15m11s&lt;/td&gt;
&lt;td style="text-align: right"&gt;7m22s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The link now took 7m22s. Still slower than the Rust build, but enough of an
improvement that I wanted to look beyond the linker.&lt;/p&gt;
&lt;h4 id="what-about-the-rest-of-the-build"&gt;What about the rest of the build?&lt;a class="heading-anchor" href="#what-about-the-rest-of-the-build" aria-label="Permalink to What about the rest of the build?"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;The build still took fifteen minutes, and nearly eight of those passed before
the linker even started. What was it waiting for? I went back to the original CI
trace to follow the inputs from Bun’s own code.&lt;/p&gt;
&lt;p&gt;buildprof also records which files each process reads and writes. If a process
reads a file another wrote, it links the two together under the hood. Turning on
“Show on timeline” draws those links as arrows. Here, the linker reads
&lt;code&gt;libbun-profile.a&lt;/code&gt; from the C++ compilation and &lt;code&gt;bun-zig.o&lt;/code&gt; from Zig. Both
arrive through copy steps; following those back takes us to the processes which
produced them:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-ci-dependency-combined-notes.png" alt="Following the linker’s dependency arrows through the copy steps to the C++ and Zig producers. The producer panels use the same time scale; C++ finishes first."&gt;&lt;/p&gt;
&lt;p&gt;The C++ side of the compilation finished first. The linker was waiting for
&lt;code&gt;bun-zig.o&lt;/code&gt;, so it could not begin until the Zig branch had finished too.&lt;/p&gt;
&lt;p&gt;It was at this point I went back to the Rust build and compared against how it
worked, and the main reason the Rust build was faster became obvious: Bun has
been split into &amp;gt;90 crates, while in Zig it was all trying to compile as a
single Zig module!&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/rust-crates-vs-zig-object-notes.png" alt="Cargo fanning out into named rustc processes across Bun’s crates, next to the single zig build-obj process which spawns nothing at all."&gt;&lt;/p&gt;
&lt;p&gt;This meant that the Zig build cannot parallelise the same way Rust can. I also
suspect, though I did not prove this, that it explains the slow linking: the
linker has to optimize one huge ThinLTO bitcode module instead of the same work
spread across crates.&lt;/p&gt;
&lt;p&gt;It was at this point I had to stop: to go any further, I would have to split up
the Zig module myself, and given that this code is all obsolete anyway, I didn&amp;rsquo;t
think it was worth doing that.&lt;/p&gt;
&lt;p&gt;Summarizing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The huge outlier in the initial Zig build vs the Rust build was the massive
linker step which ran alone at the end of the build.&lt;/li&gt;
&lt;li&gt;Changing the LTO settings for just Bun was not sufficient as WebKit, a
significant part of the build, still used Full LTO.&lt;/li&gt;
&lt;li&gt;Once I had done this, the Zig build dropped from twenty-four minutes to
fifteen.&lt;/li&gt;
&lt;li&gt;Even after this, linking still took 7 minutes and the whole build 15 minutes.&lt;/li&gt;
&lt;li&gt;The overwhelming difference which remained was structural: Rust spreads
compilation across &amp;gt;90 crates while the Zig build funnelled everything through
a single module.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And fwiw, the traces had also turned up a few things I couldn’t resist poking
at&amp;hellip;&lt;/p&gt;
&lt;h3 id="other-things-hiding-in-the-build"&gt;Other things hiding in the build&lt;a class="heading-anchor" href="#other-things-hiding-in-the-build" aria-label="Permalink to Other things hiding in the build"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;h4 id="a-build-can-contain-almost-anything"&gt;A build can contain almost anything&lt;a class="heading-anchor" href="#a-build-can-contain-almost-anything" aria-label="Permalink to A build can contain almost anything"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;In the middle of Bun’s CI build, I found commands asking the public internet for
the machine’s IP address, inspecting running Docker containers and reading the
latest Git commit message.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-ci-diagnostic-probes-notes.png" alt="Small CI setup commands visible in the process tree"&gt;&lt;/p&gt;
&lt;p&gt;These take well under a second altogether. Nothing to optimize but I just wasn’t
expecting to find them in a build trace.&lt;/p&gt;
&lt;h4 id="a-cold-dependency-fetch"&gt;A cold dependency fetch&lt;a class="heading-anchor" href="#a-cold-dependency-fetch" aria-label="Permalink to A cold dependency fetch"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;The builds above reused downloaded dependencies, so I also
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/scripts/capture-bun-cold-prebuilt.sh"&gt;recorded a fresh WebKit fetch&lt;/a&gt;.
Downloading and extracting the archive took about twenty seconds. For the first
twelve, all we see is Node running. Then it launches &lt;code&gt;tar&lt;/code&gt; and &lt;code&gt;gzip&lt;/code&gt;, and we
can see the extraction separately.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-cold-webkit-download-article.png" alt="A cold WebKit download and extraction"&gt;&lt;/p&gt;
&lt;h4 id="looking-inside-one-c-compilation"&gt;Looking inside one C++ compilation&lt;a class="heading-anchor" href="#looking-inside-one-c-compilation" aria-label="Permalink to Looking inside one C&amp;#43;&amp;#43; compilation"&gt;#&lt;/a&gt;
&lt;/h4&gt;
&lt;p&gt;Earlier, we followed the linker’s inputs back to Bun’s C++ compilation. We can
look inside those compiler invocations too. I picked one of the last files to
finish, &lt;code&gt;ZigGeneratedClasses.cpp&lt;/code&gt;, and
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/scripts/capture-bun-compiler-detail.sh"&gt;replayed its Ninja command&lt;/a&gt;
with &lt;code&gt;--compiler-traces&lt;/code&gt;. For Clang, buildprof enables &lt;code&gt;-ftime-trace&lt;/code&gt; and adds
its internal timings to the process timeline.&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote-ref" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="https://lalitm.com/assets/buildprof/zig-generated-classes-compiler-phases-article.png" alt="Clang’s frontend and backend phases while compiling ZigGeneratedClasses.cpp"&gt;&lt;/p&gt;
&lt;p&gt;The replay took about twelve seconds, split almost evenly between Clang’s
frontend and backend. Zooming in further, we see &lt;code&gt;ModuleInlinerWrapperPass&lt;/code&gt;, one
of the phases of Clang, accounts for over four seconds of the backend’s work.&lt;/p&gt;
&lt;h2 id="how-buildprof-works-under-the-hood"&gt;How buildprof works under the hood&lt;a class="heading-anchor" href="#how-buildprof-works-under-the-hood" aria-label="Permalink to How buildprof works under the hood"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The recording side of buildprof uses &lt;code&gt;ptrace&lt;/code&gt;, the same Linux interface used by
debuggers. I did consider both eBPF and ftrace, but &lt;code&gt;ptrace&lt;/code&gt; is just straight up
perfect for exactly this type of problem; eBPF tracing means &lt;code&gt;CAP_BPF&lt;/code&gt; and
&lt;code&gt;CAP_PERFMON&lt;/code&gt; permissions and hooking into potentially unstable
tracepoints/kernel functions. While with ftrace, I’d have to juggle tracing
instances to avoid interfering with other users, and getting the filters perfect
for just the build process and all its descendants is
cumbersome.&lt;sup id="fnref:8"&gt;&lt;a href="#fn:8" class="footnote-ref" role="doc-noteref"&gt;8&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;With &lt;code&gt;ptrace&lt;/code&gt;, I can launch the build and follow its children directly. Its
built-in events tell buildprof when processes fork, exec a new program or exit.
And for filesystem activity, buildprof uses a seccomp filter to intercept only
the calls it needs.&lt;/p&gt;
&lt;p&gt;How much buildprof costs is almost entirely down to how many files the build
opens. For ripgrep, recording barely changed the build time. Redis opened files
much more often, and recording added about five seconds:&lt;sup id="fnref:9"&gt;&lt;a href="#fn:9" class="footnote-ref" role="doc-noteref"&gt;9&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Build&lt;/th&gt;
&lt;th style="text-align: right"&gt;Untraced&lt;/th&gt;
&lt;th style="text-align: right"&gt;Processes only&lt;/th&gt;
&lt;th style="text-align: right"&gt;Processes + files&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ripgrep / Cargo&lt;/td&gt;
&lt;td style="text-align: right"&gt;12.27s&lt;/td&gt;
&lt;td style="text-align: right"&gt;12.30s&lt;/td&gt;
&lt;td style="text-align: right"&gt;12.43s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Redis / Make&lt;/td&gt;
&lt;td style="text-align: right"&gt;26.78s&lt;/td&gt;
&lt;td style="text-align: right"&gt;27.04s&lt;/td&gt;
&lt;td style="text-align: right"&gt;31.89s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If that overhead gets in the way, you can turn off filesystem tracing with
&lt;code&gt;--no-file-events&lt;/code&gt; and keep the process timeline.&lt;/p&gt;
&lt;p&gt;I work on &lt;a href="https://github.com/google/perfetto"&gt;Perfetto&lt;/a&gt;, so it was a natural
starting point for the UI; buildprof’s UI is a soft fork of the Perfetto UI. I
could have just opened the recordings on
&lt;a href="https://ui.perfetto.dev"&gt;ui.perfetto.dev&lt;/a&gt;, but I wanted control over how the
process tree was laid out, which details appeared when you clicked a command,
and things like those on-demand arrows between file producers and consumers.&lt;/p&gt;
&lt;p&gt;Fortunately, we&amp;rsquo;ve spent the last several years working on making the Perfetto
UI extensible through
&lt;a href="https://perfetto.dev/docs/contributing/ui-plugins"&gt;plugins&lt;/a&gt;. Most of
buildprof’s UI is reusing that infrastructure. Perfetto handles the hard stuff
(parsing traces, querying events, rendering the timeline and managing
workspaces) and I get to focus on what makes those things useful for builds.&lt;/p&gt;
&lt;p&gt;I plan on going into a lot more detail about the recorder and UI in a separate
technical post. &lt;a href="https://lalitm.com/page/subscribe/"&gt;Subscribe&lt;/a&gt; if you’d like to
be notified when it comes out! :)&lt;/p&gt;
&lt;h2 id="did-i-need-to-build-something-new"&gt;Did I need to build something new?&lt;a class="heading-anchor" href="#did-i-need-to-build-something-new" aria-label="Permalink to Did I need to build something new?"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;These days it&amp;rsquo;s very easy to make a tool just because you can. But that wasn’t
the case here; before building buildprof, I looked long and hard for an existing
tool that could give me this view.&lt;/p&gt;
&lt;p&gt;I started with &lt;a href="https://github.com/nico/ninjatracing"&gt;ninjatracing&lt;/a&gt;, which I’ve
used many times. It turns Ninja’s build log into a timeline showing what ran and
how much ran in parallel.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://ui.perfetto.dev/#!/?url=https%3A%2F%2Fraw.githubusercontent.com%2FLalitMaganti%2Fblog-code%2Fmain%2Fbuildprof-bun%2Fresults%2Fbun-zig-ci-ninjatracing.json"&gt;Here’s the Ninja log from the Zig-era build&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;But Ninja only sees part of Bun’s build. The scripts which invoke it are missing
from its log, and commands it runs appear as single blocks even when they launch
whole trees of subprocesses.&lt;/p&gt;
&lt;p&gt;There were several other tools, each covering different parts of the problem:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://doc.rust-lang.org/cargo/reference/timings.html"&gt;Cargo timings&lt;/a&gt; works
well for Cargo-managed builds, but cannot break down arbitrary work inside
&lt;code&gt;build.rs&lt;/code&gt; or see wrapper scripts above Cargo. In Bun, Cargo is only part of
the build:
&lt;a href="https://lalitm.com/assets/buildprof/bun-rust-ci-cargo-timings.html"&gt;the report I captured&lt;/a&gt;
covered 1m51s of a 5m40s CI build.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://clang.llvm.org/docs/ClangCommandLineReference.html#cmdoption-clang-ftime-trace"&gt;Clang’s &lt;code&gt;-ftime-trace&lt;/code&gt;&lt;/a&gt;
gave us the detail inside a compiler invocation, but cannot show what the rest
of the build is doing while
&lt;a href="https://ziglang.org/learn/overview/#performance-and-safety-choose-two"&gt;Zig&amp;rsquo;s Tracy integration&lt;/a&gt;
goes deeper still and is intended more for understanding the compiler itself.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://strace.io/"&gt;&lt;code&gt;strace&lt;/code&gt;&lt;/a&gt; and
&lt;a href="https://github.com/kxxt/tracexec"&gt;&lt;code&gt;tracexec&lt;/code&gt;&lt;/a&gt; can follow arbitrary processes
through &lt;code&gt;fork&lt;/code&gt; and &lt;code&gt;exec&lt;/code&gt;, but show general process events rather than a
build-oriented timeline.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://danielchasehooper.com/posts/syscall-build-snooping/"&gt;What the Fork&lt;/a&gt;
(&lt;a href="https://news.ycombinator.com/item?id=44902127"&gt;via&lt;/a&gt;) came closest: it follows
processes across build systems and presents a build-specific view. But as far as
I could tell, it still appears to be in private beta and there don&amp;rsquo;t seem to be
any plans to make it open source.&lt;/p&gt;
&lt;h2 id="whats-next-for-buildprof"&gt;What’s next for buildprof&lt;a class="heading-anchor" href="#whats-next-for-buildprof" aria-label="Permalink to What’s next for buildprof"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;buildprof already does what I wanted it to do, and I plan to keep working on it
as I use it on my own builds. But there are a few things I’d like to improve.&lt;/p&gt;
&lt;p&gt;Recording overhead is one; the Redis measurements showed there’s room to improve
filesystem tracing, especially for builds which open lots of files. I’d also
like to support &lt;a href="https://github.com/LalitMaganti/buildprof/issues/2"&gt;macOS&lt;/a&gt;
where I do some of my work and maybe
&lt;a href="https://github.com/LalitMaganti/buildprof/issues/3"&gt;Windows&lt;/a&gt; if there&amp;rsquo;s
interest.&lt;/p&gt;
&lt;p&gt;There are also more build systems and toolchains I’d like to test, including
npm, Gradle and Bazel. Computing critical paths would also be a big improvement:
we followed dependencies by hand in this post, but buildprof could help identify
the chain of work holding up the build and automatically annotate it.&lt;/p&gt;
&lt;p&gt;I’ll probably tackle these as and when I need them. But if you try buildprof and
there’s something you wish it could do, I’d be interested to
&lt;a href="https://github.com/LalitMaganti/buildprof/issues"&gt;hear about it&lt;/a&gt;. What people
find useful will help me decide where to spend more time.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;a class="heading-anchor" href="#conclusion" aria-label="Permalink to Conclusion"&gt;#&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;I managed to satiate my curiosity, though I ended up spending rather more time
on this than I expected. Along the way I built a tool I now want to have around
whenever a build is taking too long.&lt;/p&gt;
&lt;p&gt;I know I’ll come back to buildprof the next time a slow build annoys me. If you
have one of those builds too,
&lt;a href="https://github.com/LalitMaganti/buildprof#quick-start"&gt;give it a try&lt;/a&gt;. I’d love
to hear what you find!&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;In C and C++, a compilation unit is usually
&lt;a href="https://eel.is/c++draft/lex.separate"&gt;a source file together with its included headers&lt;/a&gt;.
Rust compiles
&lt;a href="https://doc.rust-lang.org/reference/crates-and-source-files.html"&gt;crates&lt;/a&gt;,
which can be split into
&lt;a href="https://doc.rust-lang.org/rustc/codegen-options/#codegen-units"&gt;multiple code-generation units&lt;/a&gt;.
Zig normally compiles a program’s Zig sources together as
&lt;a href="https://kristoff.it/blog/zig-new-relationship-llvm/#speeding-up-compilation"&gt;a single compilation unit&lt;/a&gt;.
Bun’s Zig compiler fork supports splitting that into multiple LLVM modules,
but its CI build
&lt;a href="https://github.com/oven-sh/bun/blob/0d9b296af33f2b851fcbf4df3e9ec89751734ba4/scripts/build/zig.ts#L34-L62"&gt;explicitly selected one when LTO was enabled&lt;/a&gt;.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;The
&lt;a href="https://github.com/oven-sh/bun/blob/bun-v1.3.14/.buildkite/ci.mjs#L555-L609"&gt;Zig-era CI build&lt;/a&gt;
ran its C++ and Zig compilation stages on separate
&lt;a href="https://buildkite.com/home/"&gt;Buildkite&lt;/a&gt; machines and passed their outputs
to a final linking stage. My script ran those stages concurrently on one
machine, waited for both outputs, copied them locally instead of
transferring them over the network, then linked them. This should preserve
the dependency graph, but due to the hardware differences and running both
stages on one machine, resource contention would obviously be quite
different. Also note that my timings are individual runs (albeit ones which
were quite stable) while Bun’s reported figures are medians.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;A process can do substantial work internally, including running multiple
threads, without launching anything else. The process timeline won’t show
that parallelism. To see inside a process, we need tracing from the program
itself, as Clang and LLD provide in the examples below.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;LLVM emits &lt;code&gt;OptModule&lt;/code&gt; from its legacy pass manager, which LLD uses for code
generation. The inlining and other IR optimization passes can appear before
it, so this bar is not the total time spent optimizing a module.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:5"&gt;
&lt;p&gt;In the earlier Full-LTO linker replay, 26,825 &lt;code&gt;OptFunction&lt;/code&gt; events with JSC
symbols total about 209s. This is summed event time, not a measurement of
JavaScriptCore’s entire contribution to the link. One
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/results/compiler-events/zig-full-lto-javascriptcore-event.png"&gt;example event&lt;/a&gt;
takes 2.94s; its symbol demangles to &lt;code&gt;JSC::JITThunks::initialize(JSC::VM&amp;amp;)&lt;/code&gt;.&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:6"&gt;
&lt;p&gt;&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/experiments/webkit-thinlto/record-comparison.sh"&gt;Recording script&lt;/a&gt;.
These timings are just for building Bun with the libraries already
available; the WebKit and ICU rebuild happened beforehand and isn’t
included. Of course, I could point buildprof at that build too, but that’s
another rabbit hole&amp;hellip; I did not rebuild a matching Full-LTO WebKit archive
as a control, so I cannot attribute every second saved to the LTO setting
alone.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:7"&gt;
&lt;p&gt;buildprof currently supports compiler traces from Clang, LLD and nightly
Rust.&amp;#160;&lt;a href="#fnref:7" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:8"&gt;
&lt;p&gt;eBPF tracing uses capabilities such as &lt;code&gt;CAP_BPF&lt;/code&gt; and &lt;code&gt;CAP_PERFMON&lt;/code&gt;, as
described in the
&lt;a href="https://android.googlesource.com/kernel/common/+/3135f5b73592988af0eb1b11ccbb72a8667be201/include/uapi/linux/capability.h"&gt;kernel’s capability definitions&lt;/a&gt;.
ftrace provides
&lt;a href="https://www.kernel.org/doc/html/latest/trace/ftrace.html"&gt;separate tracing instances and PID filters&lt;/a&gt;,
but these still need configuring and access to tracefs. &lt;code&gt;ptrace&lt;/code&gt; also
depends on the host’s security settings; containers may need additional
permissions to allow tracing child processes.&amp;#160;&lt;a href="#fnref:8" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:9"&gt;
&lt;p&gt;Medians of five clean builds per mode on the same VM, with six build jobs.
&lt;a href="https://github.com/LalitMaganti/blog-code/blob/main/buildprof-bun/results/overhead/measure-options.py"&gt;Measurement script&lt;/a&gt;.&amp;#160;&lt;a href="#fnref:9" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title>Changing Devtools Is Cheap. Owning Them Isn’t.</title><link>https://lalitm.com/post/changing-devtools-is-cheap-owning-them-isnt/</link><pubDate>Sun, 09 Aug 2026 14:53:00 +0100</pubDate><guid>https://lalitm.com/post/changing-devtools-is-cheap-owning-them-isnt/</guid><description>In Devtools must be open source1 (via), David Crawshaw makes the case that, because of coding agents, we’re now in an era where devtools will be personalized by individual users. Specifically, agents’ ability to jump into new codebases and build whatever we want means we’ll be hacking on the source of the devtools we use day to day (even those without extension APIs) adding features and automatically rebasing our patches across releases.
The argument is seductive, especially to a reader who thinks of themselves as a maker or tinkerer: after all, the idea that you can hyper-tune everything you use sounds like a utopia; it means things can work exactly how you want them to.
But I’d argue that Crawshaw underappreciates the ongoing cost when he writes
“Both the upfront fixed costs and the ongoing costs of personalizing software have disappeared.”
While AI has made the upfront cost of changing software a lot lower, properly personalizing software still requires your attention. And attention in the AI age is scarcer than ever.</description><content:encoded>&lt;p&gt;In
&lt;a href="https://blog.exe.dev/devtools-must-be-open-source"&gt;Devtools must be open source&lt;/a&gt;&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;
(&lt;a href="https://news.ycombinator.com/item?id=49156111"&gt;via&lt;/a&gt;), David Crawshaw makes the
case that, because of coding agents, we&amp;rsquo;re now in an era where devtools will be
personalized by individual users. Specifically, agents&amp;rsquo; ability to jump into new
codebases and build whatever we want means we&amp;rsquo;ll be hacking on the source of the
devtools we use day to day (even those without extension APIs) adding features
and automatically rebasing our patches across releases.&lt;/p&gt;
&lt;p&gt;The argument is seductive, especially to a reader who thinks of themselves as a
maker or tinkerer: after all, the idea that you can hyper-tune everything you
use sounds like a utopia; it means things can work &lt;em&gt;exactly&lt;/em&gt; how you want them
to.&lt;/p&gt;
&lt;p&gt;But I&amp;rsquo;d argue that Crawshaw underappreciates the ongoing cost when he writes&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Both the upfront fixed costs and the ongoing costs of personalizing software
have disappeared.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;While AI has made the upfront cost of &lt;em&gt;changing&lt;/em&gt; software a lot lower, properly
personalizing software still requires your attention. And attention in the AI
age is scarcer than ever.&lt;/p&gt;
&lt;p&gt;Having maintained an open-source devtool designed to be modified and forked for
nine years now&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;, I can say that most users don&amp;rsquo;t &lt;em&gt;want&lt;/em&gt; to customize their
devtools. They want someone else to make the tool reliable and coherent, so they
can focus on the problems they opened it to solve. They reach for source
modification only as a last resort, when a change is critical to their workflow
and no other route works.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;This is not to say that this sort of personalization won&amp;rsquo;t become more common: I
absolutely think it will. I just think it will take the form of strong core
systems with well-defined boundaries and extension points.&lt;/p&gt;
&lt;h3 id="personalization-still-needs-a-person"&gt;Personalization still needs a person&lt;a class="heading-anchor" href="#personalization-still-needs-a-person" aria-label="Permalink to Personalization still needs a person"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;As a thought experiment, imagine an open-source diff viewer with no extension
API. You find most diffs noisy, so you ask an agent to add a &amp;ldquo;focus mode&amp;rdquo; that
collapses imports, generated files, and other changes you consider mechanical.
It works well and becomes part of your normal workflow.&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;At first, life is good: everything works, and you&amp;rsquo;ve solved your problem. Then
upstream releases a new version that refactors the code you changed. As Crawshaw
suggests, you&amp;rsquo;re clever, so you&amp;rsquo;ve set up a bot to automatically rebase your
changes onto each update. It resolves any merge conflicts and moves your code to
the right place.&lt;/p&gt;
&lt;p&gt;But now suppose a few months pass and upstream makes a more substantial change:
it adds syntax-aware move detection. If a function moves between files, the
viewer now shows it as a move instead of one large deletion and addition. The
agent muddles through, rebases your focus-mode patch, and gets everything
compiling without any merge conflicts.&lt;/p&gt;
&lt;p&gt;But now what should focus mode do if the function has mostly moved but also
contains a few meaningful edits? Does it hide the whole block as a mechanical
move? Does it show only the edited lines without any surrounding context? Or
does it show the whole function?&lt;/p&gt;
&lt;p&gt;There isn&amp;rsquo;t an obviously correct answer; it depends on what &lt;em&gt;you&lt;/em&gt; want to see in
the diff. So what, are you going to interrupt your day to make this decision?&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s a central paradox here: if you&amp;rsquo;re okay with &amp;ldquo;let the agent decide&amp;rdquo;, then
you&amp;rsquo;ve delegated your authority to the agent. For small choices, that may be
perfectly adequate. But if you want the tool to work exactly how you want, you
need to inspect and direct those choices. Do you really want to have opinions
about the design of a devtool you use forever?&lt;/p&gt;
&lt;p&gt;The key is &lt;em&gt;attention&lt;/em&gt;. Any one personalized tool might be unlikely to fail on a
given day, but if you do this to every devtool you use, you multiply the number
of tools that can unexpectedly demand your attention.&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt; Worse, those failures
are unpredictable: a tool might work for months and then break at the exact
moment you urgently need it. Most engineers want to use devtools to accomplish a
task; they don&amp;rsquo;t want their attention diverted to designing and repairing them.&lt;/p&gt;
&lt;h3 id="shared-tools-need-a-shared-reality"&gt;Shared tools need a shared reality&lt;a class="heading-anchor" href="#shared-tools-need-a-shared-reality" aria-label="Permalink to Shared tools need a shared reality"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;All of the above applies to small teams as well. You can share the attention
cost, but at the end of the day, the team still has to ask, &amp;ldquo;How much time do we
want to spend on tools versus doing the actual work we&amp;rsquo;re meant to be doing?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I also want to look beyond Crawshaw&amp;rsquo;s post and consider how this would work in
larger companies: what happens when many teams independently personalize the
same shared devtool?&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve seen this firsthand: another big tech company makes extensive use of
Perfetto, and has hit this exact problem. Different teams in that company
decided to fork Perfetto and add ad hoc changes for their local needs. Now one
of the engineers there is fighting to consolidate them because of how painful it
is when every team means something different by &amp;ldquo;Perfetto&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;Imagine the same pattern with a company-wide bug tracker. Do you want every team
to use a version with subtly different meanings for status, priority,
assignment, and resolution? What happens when a bug moves between teams?
Different layouts and personal filters are harmless; the problem begins when
personalization changes the shared semantics or workflow.&lt;/p&gt;
&lt;p&gt;When a devtool mediates work between people, it also forms part of their common
language. Teaching, auditing, reproducing investigations, and verifying that
people are talking about the same thing all depend on a shared baseline.&lt;/p&gt;
&lt;h3 id="upstream-gets-more-malleable-too"&gt;Upstream gets more malleable too&lt;a class="heading-anchor" href="#upstream-gets-more-malleable-too" aria-label="Permalink to Upstream gets more malleable too"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;We should also not compare pre-AI upstream development with post-AI forks.
Maintainers can use the same agents to investigate reports, brainstorm ideas,
and prototype new features. I can certainly attest to how useful AI has been for
both implementing small feature requests from users and prototyping larger ones
to determine feasibility.&lt;/p&gt;
&lt;p&gt;In my opinion, upstream maintainers can, and should, spend the time saved on
implementation making their tools more adaptable: implementing broadly useful
features, adding configuration knobs where they make sense, and creating
extension points for recurring needs. AI lowers the cost of doing all of this,
including deciding where customization makes sense, adding more elaborate tests
on creative uses of your tools and verifying backwards compatibility as these
interfaces evolve.&lt;/p&gt;
&lt;p&gt;Upstream has a natural advantage here: any work done there benefits &lt;em&gt;everyone&lt;/em&gt;,
while a change to your personal fork benefits only you. By relying on upstream,
the attention required to build good software shifts from people who &lt;em&gt;don&amp;rsquo;t&lt;/em&gt;
want to spend it to maintainers who have chosen to care.&lt;/p&gt;
&lt;h3 id="the-building-block-economy"&gt;The building-block economy&lt;a class="heading-anchor" href="#the-building-block-economy" aria-label="Permalink to The building-block economy"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;In my opinion, there&amp;rsquo;s an alternative view that is much more likely to come
true, one described well in Mitchell Hashimoto&amp;rsquo;s article on the
&lt;a href="https://mitchellh.com/writing/building-block-economy"&gt;building-block economy&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Concretely, it accepts the same premise: agents can write lots of code and build
niche applications, tools, integrations, forks, and so on. But instead of
assuming that forks will become the norm, Hashimoto argues that high-quality,
well-documented building blocks will power this world.&lt;/p&gt;
&lt;p&gt;I tend to agree: agents are very good at composing high-quality components. If
maintainers provide those components alongside a focused application, makers can
build specialized artifacts on top while accepting the costs. This model also
creates an easy feedback loop for ideas to flow upstream because the product was
designed to be extended.&lt;/p&gt;
&lt;p&gt;I see signs that the world is already heading in this direction. For example,
&lt;a href="https://www.sawyerhood.com/blog/an-agentic-ide-that-builds-itself"&gt;&lt;code&gt;bb&lt;/code&gt;&lt;/a&gt; is a
very interesting agentic IDE that I&amp;rsquo;ve been playing around with recently. It has
a very nice experience that lets users add substantial new product surfaces
through self-modification. But the key is that those features are plugins built
around a maintained core and extension system, not changes made by forking the
project directly.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;bb&lt;/code&gt; is also only a few weeks old at the time of writing, so we cannot draw any
firm conclusions from it, but it&amp;rsquo;s an interesting sign of the future, IMO.&lt;/p&gt;
&lt;h3 id="wrapping-up"&gt;Wrapping up&lt;a class="heading-anchor" href="#wrapping-up" aria-label="Permalink to Wrapping up"&gt;#&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;I care deeply about both the world of devtools and open source, so this is
something I feel very passionate about. Having been immersed in this world for
almost a decade now, I think the future of well-built tools with thoughtful
design and well-designed extension points is bright.&lt;/p&gt;
&lt;p&gt;Sure, there will always be folks who want to fork and make ad hoc changes. These
are the same people who already maintain custom builds of their window manager
or terminal emulator, carrying a stack of patches to get everything exactly how
they want it.&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt; For them, the tinkering is part of the enjoyment and craft.&lt;/p&gt;
&lt;p&gt;But I think most users just want to get their work done with devtools, and we
owe it to them to give them a strong, dependable experience instead of asking
them to take on the burden of maintaining the product themselves.&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;While the title reflects the conclusion, IMO it&amp;rsquo;s not very reflective of
&lt;em&gt;most&lt;/em&gt; of the post, which is actually about personalization at the source
level. If you&amp;rsquo;ve read my other posts (e.g.,
&lt;a href="https://lalitm.com/perfetto-oss-company-prio/"&gt;On Perfetto, Open Source, and Company Priorities&lt;/a&gt;),
you&amp;rsquo;ll know I&amp;rsquo;m a &lt;em&gt;staunch&lt;/em&gt; believer in open source so I&amp;rsquo;m of course in full
agreement with the title and conclusion.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;I&amp;rsquo;m a co-founding engineer on
&lt;a href="https://github.com/google/perfetto"&gt;Perfetto&lt;/a&gt;.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;For example, the upstream project might reject a feature request because the
change conflicts with its product direction, or the tool might not expose an
extension point capable of supporting it. In those cases, modifying the
source may be the only practical option.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;Observant readers may note that this is not so dissimilar to Crawshaw&amp;rsquo;s own
example with &lt;a href="https://meat.dev"&gt;Meat&lt;/a&gt; :).&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:5"&gt;
&lt;p&gt;This is a very informal application of
&lt;a href="https://en.wikipedia.org/wiki/Lusser%27s_law"&gt;Lusser&amp;rsquo;s law&lt;/a&gt;, which says
that the reliability of a system composed of independent components in
series is the product of the reliability of those components.&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:6"&gt;
&lt;p&gt;The &lt;a href="https://suckless.org/"&gt;suckless&lt;/a&gt; ecosystem is an existing example of
this approach. &lt;a href="https://dwm.suckless.org/"&gt;&lt;code&gt;dwm&lt;/code&gt;&lt;/a&gt; and
&lt;a href="https://st.suckless.org/"&gt;&lt;code&gt;st&lt;/code&gt;&lt;/a&gt; are commonly customized by arbitrary
patches to their sources.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item></channel></rss>