Fabio Akita spent almost four months on a project of his own, nes-to-sms, a static recompiler for NES games meant to run on the Master System, before giving up on it. That was 349 commits across 31 active days over those nearly four months, and 668 tests. Super Mario Bros. came out playable, backed by a complete, public disassembly the community has used for more than a decade. The other games in the same batch, without that map, died a few frames after boot.

He told the story in an article published on September 28, 2026, “Os Limites das LLMs: Boas em reproduzir o que existe, caras pro que não existe” (The Limits of LLMs: good at reproducing what exists, expensive for what doesn’t). The line that sums up his argument: “the model knows deeply what the world has already published, and is blind to what the world keeps closed.” That limit is the reason AI is changing the relationship of anyone who works with open source software: it can read, adapt and extend a program whose code the model has already seen. The same mechanism that helps here is what gets stuck further on, and it is worth understanding both sides before relying on it. That leaves a loose question: if AI is good at reproducing what already exists, why is it being used, above all, to build from scratch what already exists?
What the model has already read
One of the largest open source training sets, The Stack v2, gathers 67.53 TB of code files uncompressed, pulled from the Software Heritage archive (see the dataset page). This is the public share of software: libraries, frameworks and scripts published in open repositories, the material a model can index and reproduce from memory.
On the other side sits what never reaches that training set at all. According to the GitHub Octoverse 2025 report, 81.5% of contributions in 2025 happened in private repositories. A company’s internal system, embedded code, anything never published: that is exactly where AI has nothing to reproduce, because it never saw it.
Why this favors anyone already working with open source
I have written here before about QGC4QGIS: bringing into QGIS the photogrammetric flight grid calculation that QGroundControl already did, reading the source of one open project to reimplement the function in another. The work moved fast because both programs are open: this is the kind of code that goes into training.
Akita’s argument explains why: reading, adapting and extending a published program is the task where AI reproduces what exists well. As soon as either side is closed or unprecedented, the model loses its reference.
Anyone who already lives in free software out of habit gets a head start for the same reason. When I wrote about using Linux with AI, I noted that on a free system the solution that shows up first also tends to be free: instead of buying the license, you look for the package, the script, the plugin. That is the same terrain where AI reads best.
What the 2023 numbers said, and what has changed
A preprint from February 2023, arXiv 2302.06590, measured the effect of GitHub Copilot on an isolated task: implementing an HTTP server in JavaScript as fast as possible. The group with Copilot access finished 55.8% faster than the control group. That is a real number, but from a short, well-defined task, far from the work of maintaining a system in production.
The test that came closest to that scenario arrived later, and measured the opposite. A randomized controlled trial from METR, published in July 2025 with data collected between February and June of that year, followed 16 experienced developers resolving 246 real issues in their own open source repositories, ones they already knew well. With AI (mostly Cursor Pro with Claude), they took 19% longer than without it. Before the experiment, they expected to be 24% faster; after feeling the slowdown, they still believed they had been 20% faster.
METR updated the result in February 2026. They ran a new experiment starting in August 2025, with 57 developers (10 of them from the original study), and the design had to change: a significant share refused to participate under the original conditions. Some said they would not want to do half the work without AI even for 50 dollars an hour from the study; between 30% and 50% avoided submitting tasks they did not want to face without assistance. METR itself points out that this bias cuts the sample in the opposite direction from what one might assume: the developers and tasks that would benefit most from AI tend to be left out of the experiment, so the estimate may be understating the real gain. The measured numbers do show some evidence of a speed advantage for AI: tasks 18% faster in the original group (range of -38% to +9%) and 4% faster in the new group (-15% to +9%), which METR itself calls “very weak evidence” of the real size of the effect. The team believes the productivity gain from AI is higher today than in 2025, but acknowledges it still has no way to measure by how much.
The same kind of leap applies to code evaluation. SWE-bench, published in October 2023, tests models solving 2,294 real problems drawn from issues and pull requests across 12 popular Python open source repositories. Mapped territory, exactly the kind of code that goes into training, which opens the door to contamination: the model may have seen the real fix during training and reproduced it from memory. That is why this post does not cite the benchmark’s current score. LiveCodeBench, from March 2024, was born as a response to that problem: it continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) so performance can be compared on problems published before and after each model’s training cutoff.
Where the model gets stuck
The distance between reproducing a pattern and reasoning about a new problem can be measured. GSM-Symbolic, from Apple, published in October 2024, rewrites the same school-level math problems by swapping only the numeric values, and every model tested drops in performance. Adding a clause that looks relevant but does not actually enter the calculation pushes the drop as high as 65% in the best models available. The authors raise the hypothesis: models replicate reasoning steps seen in training, rather than reasoning about the new problem. GSM1k, from Scale AI, published in May 2024, arrives at a similar result by another route: a benchmark equivalent to GSM8k but free of training contamination, where accuracy drops as much as 8%, with signs of systematic overfitting across practically every model size. ARC-AGI-2 (see the ARC Prize) follows the same line: tasks verified as easy for humans and designed to have no equivalent in training material.
When the problem is unprecedented and still worth solving, the fallback becomes brute force, and it has a price. AlphaEvolve, from DeepMind, published in May 2025 and tested on more than 50 open mathematical problems, rediscovered the state of the art in about 75% of them and improved the best known solution in 20%. The ARC Prize tested OpenAI’s o3-preview in December 2024 on ARC-AGI-1: in low-efficiency mode (172 times more compute), the model reached 87.5% on the semi-private set, at a cost of $4,560 per task, against 75.7% spending $26 per task in the cheap mode. And OpenAI claimed, according to New Scientist on September 8, 2026, to have reached a solution to the Navier-Stokes problem: a thousand agents working 50 hours on a related problem, then ten thousand agents for 11 hours to extend the result, at an estimated cost of $15 million to repeat the process on demand.
Solving an unprecedented problem is expensive even when it pays off: computational brute force charges a price proportional to the novelty. Repeating an app that already exists somewhere is cheap, and that cost difference helps explain the flood of releases that comes next.
The flood of new apps
While AI reads what is already public better, the production of new things has also accelerated, and the data on both sides date from the same recent period.
The GitHub Octoverse 2025 report, published on October 28, 2025, counted 121 million new repositories created that year (the same report behind the 81.5% of contributions in private repositories cited at the start of the post). It is a snapshot of how much new code is born, without separating an original tool from something that rewrites what already existed in another repository.
On the app side, Appfigures, cited by TechCrunch on April 18, 2026, recorded that new app launches in the first quarter of 2026 grew 60% over the same quarter of 2025, combining the App Store and Google Play, and 80% counting the App Store alone. The report calls the link to AI a working hypothesis, pointing at tools such as Claude Code and Replit, without presenting it as established fact.
The other side of the same phenomenon shows up in whoever maintains what is already open. A preprint from January 21, 2026, Vibe Coding Kills Open Source (Koren, Békés, Hinz and Lohmann), builds an economic model: when AI agents assemble software from open components without the end user ever interacting with the source project, and maintenance depends on that engagement, higher adoption reduces the inflow of new contributors and sharing, which impoverishes the availability and quality of the open ecosystem itself. It is a theoretical model, in a preprint, not a field measurement.
Concrete cases of a maintainer feeling that weight already surfaced in 2026. Adam Wathan, of Tailwind Labs, commented on a pull request on January 7, 2026 that the company laid off 75% of its engineering team because of what he called the brutal impact of AI on the business: documentation traffic fell about 40% since early 2023, with Tailwind more popular than ever, and revenue dropped close to 80%. Documentation is the only channel through which people learn about the company’s commercial products. Daniel Stenberg, maintainer of curl, announced on January 26, 2026 the end of the paid vulnerability bug bounty program, closed on January 31 that year, citing the volume of low-quality AI-generated reports. GitHub changed the product itself on June 17, 2026 to let maintainers limit how many open pull requests a contributor without write access can keep open at once.
There is no direct statistic comparing new AI-created apps against contribution to an existing project: nobody measured both ends the same way. What can be placed side by side is this: on one side, more repositories and more apps being born; on the other, maintainers of popular open projects reporting less engagement and more noise. In my reading, the two pictures converge: generating got cheap, and contributing still requires reading someone else’s project.
Before building, search
A tool that already exists carries what a new project does not yet have: people using it, cases that have already tested the software in real use, documentation written by someone who has already struggled with the same problems, and someone who took on maintaining it. An app built from scratch starts with none of that, even when it solves exactly the same problem.
It is also where AI pays off most. As the previous sections show, the model reads what has already been published well: understanding an open repository, finding where the function that matters lives, and extending or fixing what already runs is the task it gets right most often. Generating a clone is also cheap, because the model has already seen something similar; what it does not deliver along with it is the usage history, the tests, and the person who answers when something breaks.
What is left over after publishing is a bill that falls on whoever published it. I have written about this before in the section “The cost that shows up later”, in the post about QGC4QGIS: writing got cheap, maintaining did not. Every new tool, adopted or not, becomes one more project that will need to keep up with the next version of the system it runs on, and nobody does that for free forever.
Before opening a new project, three questions are worth asking: does this tool already exist somewhere? Where is its repository? What is actually missing from it that justifies not using the one already built? Sometimes the answer is to open one anyway, because what is missing is too big for a tweak. Often, what is missing is already described in an open issue, waiting for someone willing to read the code and fix it.
What I did inside QGIS
The five plugins I maintain follow the same pattern: instead of a new site or app, the functionality went inside QGIS, on top of the project the user already had open.
GisBR replaces what used to require programming in R or Python with IPEA’s geobr package, or manually browsing FTP, WFS and ArcGIS REST servers looking for official datasets. Outside QGIS, that would have become one more download script. Inside it, it became 55 algorithms in the Processing Toolbox, with local caching, automatic mirroring on GitHub, and built-in SSL certificate support for the main official Brazilian connectors (post, portfolio).
Desire Lines comes from flow maps I used to teach people to build with SQL directly in QGIS, or from old tools that simply stopped running. On its own, it would have been another script to paste into an origin-destination matrix. Inside QGIS, it is today a plugin approved in the official repository, generating desire lines from the matrix without requiring a single line of code from the user (post, portfolio).
SIG-Bus was born to end the manual matching between Belo Horizonte’s GTFS and the PBH boarding CSV, which required understanding both formats and resolving by hand the route identifiers that did not match between them. On its own, it would have been an ETL script run once per analysis. As a plugin, it imports the GTFS, validates its integrity, matches demand through a spatial join and returns schedule and ridership-by-segment layers already inside the project (post, portfolio).
Logis occupies the space between proprietary routing software, too expensive for a small municipality, and a researcher’s script, which requires a Python environment already set up. Outside QGIS it would have been one more script of that second kind. Inside it, it became 25 algorithms organized into three modules, and the urban logistics module reuses the same OpenStreetMap road network pipeline GisBR had already solved, which saved the most time-consuming part of the work (post, portfolio).
QGC4QGIS, already mentioned earlier in this post and in the text about the open ecosystem, replaces the manual replanning of the flight grid inside QGroundControl. The isolated alternative was to keep redrawing the polygon by hand in the flight app, over a generic satellite image. Inside QGIS, the grid is born over the project’s own layers, which already have the area boundary, the road network and the terrain (portfolio).
None of these five is a contribution to QGIS’s core: they are extensions, plugins that only exist because the platform already exists first. Even so, they count as cooperation, because they run on top of what QGIS had already solved before (format reading, reprojection, editing, printing, the Toolbox itself) and reach anyone who already has the program open, without asking for anything else to be installed. The limit is the same one I have written about before: for anyone who never opens QGIS, none of these five plugins is of any use.
What changes for those who use and maintain software
The thread running through this post is always the same: AI reads well what is already public and does little for the question that never had a written answer anywhere. That holds for the code it helps extend, for the app someone decides to build from scratch instead of searching first, and for the plugins I have been publishing inside QGIS instead of outside it.
Where the software is closed, the problem is unprecedented, or there simply is no published answer already out there, the nature of the work for whoever operates the agent changes. What is left is the part Akita describes as the oracle: defining what the result will be checked against, whether an automated test, a reference calculation, or a manual inspection. And what is left is recognizing the moment the agent stopped converging and started circling the same error, reverting one attempt after another without moving forward, the same judgment call I covered in the post about learning with AI without autopilot. Neither of these can be delegated to the model itself.
Before asking the agent to build something, it is worth asking two things in the right order: does that already exist somewhere, and where can I contribute to what already exists, before asking how to build it again.