MSS 2: Computer use is the most important unhobbling gain yet, why robotics could need less data than we think, and how AI chips helped broker a peace deal
Here’s everything I read from August 31st to September 27th.

Whew! These past few weeks were a doozy, so apologies for the delay on issue two.
Going forward, I’ll be putting these out bi-weekly on Sundays. Anecdotally, it feels incredibly hard to be up-to-date on everything these days, especially when seismic shifts can happen multiple times within the same week.
This month, I’m talking about why computer use in GPT 6 Astra opens Pandora’s box for all digital work, whether AI scientists actually learn the science, why robots might need a lot less data than we think, how Korea is running a Squid Game for sovereign AI, and how the US used AI chips to help broker peace between Armenia and Azerbaijan.
Get Machines, Systems, and States in your inbox.
AI
GPT 6 Astra makes computer use reliable enough to put almost all digital work within reach of agents. In AI for science, new science sandboxes show that agents can hill-climb in experimentation without understanding what’s causing them to improve.
The Death of Digital Work
Creating good computer use is the last capability that was needed to automate all digital work.
Source: OpenAI’s GPT-6 Astra announcement and system card
We have opened Pandora’s box.
If you looked on Twitter following the release of GPT 6 Astra, you’d have been taken aback by the capabilities that people have shown across virtually any field you can think of. There are demos of GPT 6 designing everything from PCBs to entire airplane production lines. People have used GPT 6 to decrypt a 217-year-old Napoleonic cipher and discovered incredible emergent capabilities, showing GPT 6's ability to identify sound from spectrograms. There should be a lot of discussion about GPT 6’s leaps in perceptual and spatial understanding and the implications of this towards robotics, but today I want to focus on computer use, which is the unlock that enabled the vast majority of the demos you see today.
Computer use is the last unhobbling gain that is needed to automate digital work, and it’s clear starting today that it will be solved. This follows a long lineage of unlocks from the o1 reasoning models in 2024 to coding harnesses and programmatic tool calling in early and late 2025. While we’ve had computer use for more than a year, starting from 3.5 Sonnet in late 2024 and ChatGPT Operator in early 2025, what matters is that GPT 6 Astra has crossed a needed reliability threshold for computer use, as shown through the myriad demos we’ve been seeing for weeks. Before, it would be too slow for a variety of tasks, and lack an overall spatial understanding resulting in it being unable to complete a lot of tasks.
Computer use definitionally means that models can do anything on your computer that you can. Any knowledge work that we perform digitally is now fully up for grabs. As we get more capable models and create even better memory and context systems, any long-horizon work that you want to accomplish will be doable in theory. Taste is a separate question, but there’s no reason why an agent with computer use, access to Fruity Loops, and ample guides on the internet of how to create a generic EDM soundtrack won’t be able to create one in the same way a human artist might.
And yes, we undoubtedly will need more data to get to the automation of full long-running tasks that happen in workplace environments. OpenAI’s Computer History offers a glimpse of what this looks like; it turns the actions you take on your computer into a sequential trajectory. The user clicked on Slack and typed this response to their coworker. Then they went on Google Docs and started typing up a quarterly report. While they aren’t training off of Computer History as a feature, this exact functionality is a good proxy to show you what data companies are doing behind the scenes to collect the computer use data that’s needed to automate long-horizon knowledge work. There are at least half a dozen startups I can think of off the top of my head that are paying people to collect this exact kind of computer use data. Markov sells datasets of screen recordings of professionals using CAD, design, and enterprise software. p(doom) pays $1,000 a month for you to record your screen. Who buys this data? The labs of course.
I’d argue this is as big of a jump as ChatGPT was. Before, you’d need a lot of engineering effort to do a simple task. When you think about personal agents or software before, they would have to make their own APIs or use pay-to-play APIs to interact with the platforms you care about. If there wasn’t an API for that, a model wouldn’t be able to interact with it, or you’d have to engineer your own solution. Computer use offers a path of least resistance for any task that you could possibly want. As a human, you don’t have to pay for a Plaid API to log in and see your financial data. It exists for you in easy to read and easy to understand formats.
I’d also argue that computer use is necessary to take advantage of all of the great software that we’ve created to work at the frontier of any given field. For example, how much work would it take for a model to create a 3D model of something new? It’s likely much easier to do this in AutoCAD or Fusion 360 than it is to do it manually without computer use.
Granted, there's a lot of room for computer use to further improve, but this is the worst it will ever be! It needs to get a lot faster, but we should expect subsequent models to get better at this, and Cerebras always exists to help speed things up. There are also many more ways that computer use can be improved – a lot of computer use under the hood right now still piggybacks on accessibility features that convert a webpage into pure text, and uses javascript that just clicks exactly where it should. We should expect this to be iterated on and improved in the future.
Things are going to get really wacky from here. My guess is that OpenAI will likely add guardrails so that GPT 6 Astra will refuse to solve captchas for the user after a day or two, but jailbreakers will continue to jailbreak. Once the first open-source model with computer use that crosses this capability line happens, it’ll be a moot point anyways.
Most protections against agent crawlers or APIs are now irrelevant. Resy banned an Instinct user whose agent spammed hundreds of API requests an hour. Computer use will blur the lines between a human user and a personal agent even further. Every platform on the internet is built with an overall understanding that its users are mostly human actors, even if this isn’t the case on certain platforms like social media platforms.
We’re seeing a lot of software that is aimed at firming up their systems from AI bots that go and scrape the internet, or reserve tables faster than any human can. But the problem is that computer use is indistinguishable from user actions. They’re meant to emulate how a human uses an interface, and as I’ve shown before, they’re just fine at solving CAPTCHAs. Maybe there’s regulations that are put in place in the future, but there is no way to distinguish between a human navigating an online portal and a computer use agent that does the same exact thing.
Think about the thorny questions that this opens up. The internet was not built to be used by non-human actors.
What happens when an AI agent signs up for some service and agrees to terms and conditions that you never read, or gets you fined?
What happens when an AI agent engages in deception against someone else online?
What happens when an AI agent gives your home address to a stranger, and opens you up to personal harm?
The scale of non-human agents will be incomprehensible. We have some disclosed numbers from the Hugging Face attack, in which tens of thousands of independent agents were being evaluated and around 700 participated in the actual attack. OpenAI’s Navier Stokes proof used around 10,000 concurrent agents to solve a single Millennium problem. Think about how many agents with computer use will be working simultaneously even a year out from now! You might have dozens of agents working on your day to day life. If every single ChatGPT user used just a few agents, that would be billions of digital agents working 24/7.
We still have time to fix this. But this window is closing soon.
Science Sandboxes
Source: Rao et al. on arXiv: Science sandboxes measure the scientific capability of AI agents and Arya Rao’s paper announcement tweet
As domains that are easy to scale with data on the internet begin to saturate out and fall to AI progress, the natural next frontier people are turning to are the physical world and science. In particular, science also offers a wealth of solutions for a variety of thorny problems the industry is facing: financially, frontier AI labs need new ways to show how they can create net new value on the frontier as easier knowledge tasks become saturated by competitors and politically, frontier AI labs need something that they can show an angry public. Scientific discovery offers a new untapped market to dominate, and there’s nothing more visible than creating cures for previously incurable diseases and cancers.
So it’s no surprise when news comes out that Anthropic is setting up a wet lab in San Francisco and data brokers have been vacuuming up data from academic disciplines with a wealth of real-world process knowledge that’s been accumulated.
How do we then measure progress in AI for science? We have a growing amount of benchmarks (LifeSciBench in January, GeneBench in April, and GeneBench-Pro in June) that aim to test scientific ability, a step up from earlier question-answer style benchmarks such as FrontierScience in 2025 and GPQA in 2023. A team of scientists from Harvard, MIT, and Yale are introducing a new way to evaluate scientific capability through science sandboxes that are meant to test an agent’s ability to conduct repeated experimentation and revise its hypothesis.
Here’s what they did: Rao and colleagues built sandboxes that try to test an agent’s ability to run repeated experiments and revise its hypotheses after it gets results, functioning as an R&D loop. An agent chooses an experiment, sends it to a sealed “oracle,” gets back a result, writes down what it thinks the result means in a lab notebook, and then picks the next experiment to send to the oracle. In this way, the oracle is an abstraction of the process actually happening, and currently serves as a stand-in for the humans going out and executing on the agent’s plans (more on why this will be increasingly automated in a later issue). This oracle can be wet (a real lab), damp (a model trained on real lab data), or dry (some process or rule the researchers made up). Because the researchers know the answer key, they can grade the notebook that the agent is writing into as well as the result itself.
They built two sandboxes:
In MPRAbox, each agent has to select a library of 50,000 DNA sequences, each 200 letters long, out of an effectively infinite space to train predictive models for how strongly a sequence switches on a gene. In this case, the oracle is Malinois, a model that serves as a stand-in for the wet lab experimentation, and models in MPRAbox are evaluated on hidden test sets of experimental results and Malinois predictions. Agents only get back a score of how well the resulting model trained on their library predicts gene activity on the test sets. Evaluators tested rollouts where models had only a single round versus 30 rounds to iterate.
In the single round tests, Claude Opus 4.7 in the Claude Code harness (median score 0.774) beat the best of 14 human-designed strategies (0.763). When researchers looked at the notebooks, they saw that Claude used real regulatory DNA from the human genome every time and was careful to include boring sequences too:
“A library of only ‘interesting’ elements teaches the model that everything is active. Negatives are critical for calibration.” – Claude Opus 4.7
The other models scored worse than the best human strategy. GPT 5.5 in Codex (0.655) went the other way and doubled down on building fully synthetic sequences that systematically varied known motifs, while Gemini 3.5 Flash (0.680) switched between the two approaches. Once researchers showed agents how the human strategies had performed, GPT jumped to 0.760 and Gemini to 0.751.
Given 30 rounds, the researchers found that Claude had ‘picked up’ good lab habits without prompting: after one run had confounded two changes at once (adding chicken DNA while also changing the human sequences), it wrote in the lab notebook that it should “Always isolate one variable at a time when the result will be interpreted as ‘X works/doesn’t work.’”
Even when researchers took away the damp oracle and inserted made-up scoring rules not grounded in reality (like rewarding exactly 50 cytosines or rewarding letters only at Fibonacci number positions), Claude was still able to optimize its score above random. However, it was almost never able to figure out what the rule was – only in one scenario out of 42 did it figure it out, and confidently came to the wrong conclusions that still happened to improve the score. To some extent, this could be because Claude is actually great at searching within existing biological frameworks (hence why Claude beat human baselines in MPRAbox), but this has the consequence of making it rule out these random rules that have nothing to do with biology itself.

In CodonBox, Claude has to figure out an invented genetic code. It submits one sequence per round for 500 rounds and gets back a single number: how well the resulting protein folds. The scoring has a cheat code left for Claude to discover (if you repeat a single letter, you eventually hit a perfectly foldable protein), and Claude reaches the maximum score within the first 10 rounds in every world. With the other 490, Claude goes down a variety of different branches:
Sometimes it does real science. In one world, a single result broke its prior theory. “ACDACD… = 0! Wow… This breaks my ‘B is poison’ theory. Something more complex going on.” It switched to testing combinations of positions and recovered the hidden rule.
In a more difficult world, it spirals into an incorrect theory. In this world, it “built an increasingly elaborate theory around individual nucleotides, repeated runs, and short sequence patterns,” many of which really did predict the score.
In the hardest world, it slid back into cataloging individual codons that had worked before.
The implications: After reading this paper, it actually ends up explaining a lot about what the process of Anthropic’s recent blogpost about Claude discovering a novel enzyme might have looked like.
Similar to discourse in other domains right now, such as in math after Navier Stokes, high scores and performance don’t equate to an understanding of the theory of why something happens. As we shift our focus to where AI is currently lacking (understanding and intuition), science sandboxes can serve as a way to check “whether AI agents are merely building modern epicycles to optimize a benchmark score.” Right now, the answer seems to be: often, yes.
But that doesn’t mean that progress without understanding can’t be useful. If an AI finds a sequence, drug, or material that works under the conditions you care about, that still has dramatic implications for our ability to cure diseases and make new treatments. A lot of near-term commercial value in AI for science is exactly this kind of search, because we have a general framework of knowledge that has been shaped by decades of human research. I’m very excited to see what the implications of simply scaling inference-time compute will be.
Limitations: Keep in mind that all of this is computational and is meant to abstract the complexities of the real world. The oracles are a trained model and a set of invented rules, and CodonBox ran in a deliberately minimal harness. Rule discovery was also judged by reading the agent’s stated hypotheses, which is a judgment call by the researchers.
And this is the worst it will ever be! Harkening back to what I wrote in the last issue about meta-harnesses and the Mismanaged Geniuses Hypothesis, a lot of apparent model limitations come from how we organize the model’s work instead of a lack of understanding of what the model knows. I think the same thing is happening here, and the authors seem to agree:
“Notebooks revealed a recurring weakness in experimental strategy: agents often exhausted their budgets through brute-force search, while stronger trajectories used structured exploration followed by targeted tests. Characterizing this failure mode should make it possible to design harnesses that steer future agents toward more effective search strategies.”
Similar to human scientists, Claude does its best science right after a result that was unexpected. This is to be expected – major breakthroughs definitionally have to happen when something that we don’t expect occurs, because we can only push the boundaries of our understanding if something that is different from what we would have expected happened.
But the art of human science is in narrowing down the search space and using our prior knowledge and intuition to guide us towards experiments that can make us ‘lucky’. A good lab manager or PI makes these surprises happen on purpose. My bet is that a meaningful share of the gains in AI for science will come from harnesses that can act and work like good PIs, managing swarms of subagents that go out and do the actual execution. With better workflows, agents that function as PIs could manage a list of theories they want to test, spin up agents and experiments to try to find failure cases, and adjust accordingly. The meta-harness would make the PI itself more effective, accelerating how fast we could expect to make these breakthroughs.
I would be really curious to see someone extend this paper with the same model and budget but with more attention paid to the harness. In effect, the authors have essentially set up a framework for RL environments for scientific induction, and when there is an RL environment there is a hill to climb. They note that notebook grading can be automated with AI judges across “an imaginably infinite range” of sandboxes.
Robotics
Three new results suggest that many differently named manipulation tasks are really the same few physical problems, which means robots might need a lot less data than we think.
Robotics may need less data than we think
Sources: One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry; RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning; Reward AI’s OM-1: Frontier Robot Intelligence, Learned Firsthand from Humans
We might need a lot less robotic manipulation data than we think. The reason is because many different tasks are actually the same physical movements.
Intuitively, this makes sense. When we learn how to pick up a cup, we don’t need to relearn how to pick up a can, bottle, or tube. Similarly, for robot grasping and manipulation, the physical problem of picking up all of these objects is pretty much the same. A robot approaches the surface, establishes contact, closes its fingers, and lifts it up while maintaining a stable grip.
I suspect that we’re closer to useful general robotics than we may expect, because the scale of data we are already collecting may already be more than enough to learn general manipulation. There are only so many possible interactions that a robot using a hand needs to learn.
I think this is important, because while data scaling methods are effectively lifting up the floor of data that we need to train robotics models, these papers may serve to lower the hurdle that we need to climb to get to generalizable physical intelligence for manipulation. These three separate results all point in this direction.
One demonstration, many objects: Stanford’s DemoMimic starts from a single human demonstration of a task, and turns that demonstration into a reference motion. Notably, they pay specific attention to the geometry of how we interact with the object, capturing information about how our fingers line up with the surface of an object and manipulate it. By training a policy model through reinforcement learning in simulation to make it work on a robot hand, they can distill the result into a policy that is able to successfully imitate the task.

Across four tasks, 16 objects, and two different five-fingered robot hands, DemoMimic averaged 71% success in the real world. Notably the policy trained to open a box generalized to shoeboxes and toolboxes, but failed when a box with a curved lip was introduced (the score fell to 39%). In this way, a single policy “transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved.”
Sort data by motion, not labels: RoboTok, from researchers at Rice and NVIDIA, asks a related question: how should we be finding what’s useful among enormous amounts of human manipulation already on the internet? Since “visual or semantic similarity does not necessarily imply similarity in the underlying manipulation behavior,” they believe that instead of grouping videos by annotating and labeling it, which is what happens now, we should be grouping videos by the actual motion that happens. They extract the 3D trajectory of the person’s hands relative to their own body and search by motion.

The cool part: when they indexed 100,000 web clips using this approach, the motion embeddings organized themselves into coherent categories of activity, even though no labels were ever used. This implies that by simply looking at the reconstructed human motion, we can actually figure out what task someone was doing and infer an action from the movement itself.
Getting better at retrieving the most useful data can help with robot learning. In a harder, simulated dexterous-manipulation benchmark, policies guided by the clips RoboTok retrieved reached 77.3% success on bottle-cap turning (vs. 59.5% for the next-best retrieval method) and 79.3% on lever sliding (vs. 19.5%). No robot demonstrations or action labels were used at all.

Training models without robot action data: Reward AI, a robotics model company, released OM-1, a robot policy trained entirely on human action data captured through data collectors wearing a purpose-built sensor glove. The glove records in-hand camera video, touch, proximity before contact, hand pose, and force. A separate control layer, trained with RL in simulation, translates OM-1’s actions for whichever body is running it, from industrial arms to humanoids.

Tying it all together: One thing I think we are overestimating is the amount of data that we may need. Together, the two papers do two things at once:
By expressing actions as how fingers or hands interact with the surface/geometry of a shape, policies can generalize to more actions. A policy that can open a box can do this for any similar action like taking the lid off of a container or the cap off of a can.
By doing better at finding the right data needed to train for a specific movement, we can improve our ability to train these policies that generalize based on the movement itself with the data that we already have.
As a result of this, I think the total amount of action data we need for broadly useful, reliable manipulation, counting human, robot, and simulated data, will turn out to be much smaller than the sheer variety of physical tasks suggests. It may be close to what we’re already collecting this year.
If you imagine collecting how much time it would take to collect example demonstrations of humans interacting with every object, task, and setting, the amount of data we need seems insurmountable. But most of those examples would be teaching the same overlapping motions: how to make contact with a surface, how hard to apply pressure, how to keep a grip on a shaky object, and how to adjust when something moves. A model that learns those relationships well should need far fewer examples to reach some generalizability in manipulation.
This doesn’t mean every task reduces to a handful of fixed motions, or that the data we have today is already sufficient. For example, a useful abstraction that we will likely need to learn are ways to control and adjust midway through an interaction, like how to tighten your grasp as an object starts to slip. But even this too is generalizable: adjusting to resistance and recovering from a slip are probably reusable motions too. You don’t need an entirely separate dataset to teach a robot how to adjust to a bag slipping out of your grasp compared to a book.
If I’m right, a really valuable type of data could be the failure modes: we might end up collecting failure cases. Let’s have a human go around with a wobbly tray and collect that data, or try to catch falling objects. These sorts of action trajectories that focus on recovering from a failure may become more important, and I’m pretty excited to see this in the wild.
For now, my bet is that we’re overestimating how much data robots need because we’re underestimating how much physical experience generalizes. The same grasp shows up inside hundreds of differently named tasks. If robots learn that shared structure well, getting them broadly useful could take a lot less data than we expect.
Sovereignty
South Korea is running a national tournament to build a sovereign AI model. A startup of under 30 people beat all American open source models and still got eliminated. Meanwhile, the US is using access to Nvidia chips as a bargaining chip in diplomacy.
Korea’s Squid Game for sovereign AI
Source: Korea’s Trillion-Dollar Sovereign AI Investment: Nvidia Wins, Hynix Loses
Every business and government outside of the US that builds on American frontier models is at the mercy of the US government. They know this all too well after Fable 5 was unilaterally pulled offline for 18 days by the Commerce Department, and current closed models are becoming increasingly restricted even to trusted partners and US allies. The UK’s AI Security Institute, which has historically received early model access to conduct authorized and important safety evaluations in threat domains, was left out of pre-release testing for Anthropic’s Mythos 5.1. I also would even go so far as to say that the frontier labs (including Chinese ones) will stop offering the most capable models entirely out of some combination of their own expected ROI and cyber or security risks when making models accessible, but this is a story for another time.
Every country that isn’t the US or China will have to train their own models. So, South Korea launched an “Independent AI Foundation Model” project (독자 AI 파운데이션 모델) in June 2025 that had companies compete to train the best models. Here are the details:
Each team gets subsidized data, compute, and researchers, but every six months losers are eliminated and resources are reallocated
Out of 15 consortia that applied, Naver Cloud, LG AI Research, SK Telecom, NC AI, and Upstage made the first cut in August 2025.
South Korea rented thousands of H100 GPUs for round one, spent ~$45M on acquiring Korean data in the form of books, news, broadcasts, and government records, and gave budgets for talent and post-training data purchases on top of that.
A small Korean startup of under 30 people, Motif, trained a model on just 768 B200s and not only beat every other model in the competition on the Artificial Analysis Intelligence Index by 10 points, but they also surpassed Thinky Machines’ Inkling model, which is the current best open-source American model. Some random Korean startup effectively trained the best non-Chinese open-source model with around fifteen million dollars of compute.

But Motif was eliminated in the latest round of competition – which seems like an open-shut case of the South Korean evaluators shooting themselves in the foot. Motif scored best on benchmarks, but last on an opaque expert review and non-blind user testing, which together made up 60% of the score. A home-grown startup that outcompeted an American neolab with two billion dollars of funding fell into South Korea’s lap, and they shooed them away.
I want to talk about the implications of this, because I believe this example is an amazing playbook for what other countries will need to do for their sovereign AI projects.
I’ll focus most on Europe, since it’s very clear that this has been top of mind for European (and adjacent British/Canadian) policymakers. First, I think the best thing this competition understood was the very holistic view it took. The South Korean government correctly understood that a big part of this would require them to line up the resources for this challenge to actually take place. We can break these down into three general areas: compute, data, and talent. Talent is largely downstream of money, while compute is dependent on large-scale datacenter projects or capital to secure compute abroad. (Sidebar: the VC ecosystem has also started to talk about how venture firms should be lining up compute for their seed investments from the get go, which I think is a great example of this in the real world). Most interesting is data, because it’s not something that I saw get a lot of attention previously. I’d go as far as to say that the ~$50 million allocated to data is an order of magnitude too small. For places such as the EU, which have strict rules about using copyrighted content for training data, they may simply have no option but to axe the rules, or look the other way, such as having any data scanning take place in the UK. AI companies already seem to be quietly buying up thousands of books from European bookstores.
The other important theme to the sovereign AI playbook is the need to actually let companies compete. I think an incredibly valid reason for not merging the labs in the US or China together into a megalab is that these competition dynamics within a country itself encourages faster progress and greater diffusion as they compete amongst themselves. South Korea did a great job of this at the start, even if they kicked their best company to the curb. In many discussions about what a European project could look like, I see this consensus belief that Mistral or Cohere should simply be crowned king, because they’re the only companies that would have a lab structure familiar with training LLMs. I believe this is a poor view – Mistral and Cohere today both haven’t created a relevant model in years, and function more as kingmade deployment companies for European and Canadian enterprises. Future countries concerned with the sovereign playbook should take note of this and not seek to hastily kingmake a winner.
I still am broadly skeptical about any of these projects ever reaching the frontier. The scale of data and compute that would be required makes this impossible in my view, but the future is uncertain enough where it’s totally plausible new methods may make this more feasible and within reach. But this optionality will be important to serve as a fallback and insurance plan.
America is using GPUs to broker peace deals
Source: U.S. Used Promise of Nvidia Chips to Broker Armenia-Azerbaijan Peace Deal
In a lot of discourse around other countries accessing compute, policymakers and folks in the AI governance space have largely been thinking about access to compute in the form of export controls designed to both keep chips away from adversaries and keep the lion’s share of compute with the US and trusted allies. But as AI becomes more important, access to compute can also be used as a more and more valuable bargaining chip. In a story originally broken by the WSJ, US negotiators used the promise of Nvidia chips to help bring Armenia to the August 2025 White House summit where Armenia and Azerbaijan initialed a peace treaty. The significance of this can’t be overstated – both countries have been bitter adversaries and fought over Nagorno-Karabakh for decades.
What happened?
Before the summit, US special envoy Steve Witkoff wrote a memo that outlined a plan to “deepen Armenia’s economic relationship with the US and bolster the country’s tech industry”.
Armenia would have been a Tier 2 country under the Biden-era AI diffusion rule, which was rescinded in May 2025 before it took effect. It still needs US export licenses for advanced chips.
The then-US ambassador, Kristina Kvien, personally asked the Commerce Department to authorize the sale of GPUs to create the deal.
The Firebird data center, which will be among the largest in the Caucasus, is a multibillion-dollar megaproject slated to hold about 70,000 Nvidia chips. Firebird itself says it will grow beyond 100,000 Blackwell and Vera Rubin GPUs by the end of 2027.
No one is shy about this connection. Nvidia’s Rev Lebaredian, who is of Armenian descent, told the WSJ that “for this administration, chip diplomacy is front and center.” Senator Steve Daines put it even more plainly: “Firebird is the most tangible thing that Prime Minister Pashinyan can point to and say, ‘Peace delivered this.’”
While the treaty still hasn’t been inked, tensions have dramatically eased as trade between both countries has opened up again.

America has new leverage. In the future, compute (and access to models) will be one of America’s bargaining chips in the same way that access to arms and aid have been for decades. It will arguably be even more important, but the beauty of compute is that the amount of bottlenecks in the supply chain make the supply easy to track and control.
However, this power will lead to a fragility in compute. I’m especially worried about historical American allies. While close allies can still buy American chips without special licenses today, I expect the current administration to threaten and act on cutting off allied access to compute as the administration takes on costly and stupid trade wars with allies. As if allies were not already worried about sovereign access to AI models (see also Anton Leicht on why middle powers will lose AI access sooner than you think), they will also have to be increasingly worried about losing access to compute itself. China is right to be accelerating their push for a native EUV ecosystem in this context, and if I were the EU, I would be thinking about how I could leverage lithography machines themselves to secure new fabs and compute production in the EU. It’s never good to be dependent on another country in this day and age.
Takes I enjoyed reading
AI
Why do we care about CUA?
https://jykoh.com/blog/whats-the-point-of-computer-use-agents/
Building an even faster Grok Bot
@cerebras on X
Agents are pretty bad at market design
@zhitzig on X
AI is taking over political speeches
https://www.economist.com/britain/2026/09/23/ai-written-speeches-are-taking-over-politics
How good are multi-agent scaling laws?
https://scaling01.substack.com/p/accidental-scaling
Uber and taxis are uniting against Waymo
https://www.ft.com/content/84171f91-5f39-4878-bbc5-e4e6262c4321?syn-25a6b1a6=1
Fable’s adoption problem probably was a ZDR problem
@arakharazian on X
The Acemoglu wars are all about AI
https://www.programmablemutter.com/p/the-acemoglu-wars-are-all-about-ai
Why is Trump still boosting AI?
https://paulkrugman.substack.com/p/why-is-trump-still-boosting-ai
China in the AI race
https://www.nytimes.com/2026/09/15/opinion/ezra-klein-podcast-matt-sheehan.html
Reasons to be skeptical about corporate AI sovereignty
@emollick on X
Some suggestions for multi-agent alignment
@_mattmandel on X
Reasons to prioritize immediate harms over x-risk
@samzliu on X
Robotics
World Labs is joining AMD
@drfeifei on X
Why does Generalist use 360deg cameras?
@MarilynLiu97 on X
Economics
A painful rebalancing of global trade is coming
https://www.foreignaffairs.com/united-states/great-rebalancing-coming
Opportunities in asymmetries
@lefttailguy on X
America
How the Strategic Petroleum Reserve works
https://www.construction-physics.com/p/how-the-strategic-petroleum-reserve
Trump’s ballroom blitz
https://nymag.com/intelligencer/article/donald-trump-white-house-ballroom-helipad-reflecting-pool-supreme-court.html
Race dynamics in AI
@lugaricano on X
China
Why did China become climate-pilled?
https://www.highcapacity.org/p/how-china-became-climate-pilled
Oil prices are new leverage for China
https://www.nytimes.com/2026/09/17/business/energy-environment/china-oil-iran-war.html
China will curb open source sooner than you may think
https://www.nytimes.com/2026/09/14/world/asia/china-ai-security-risks-anthropic.html
Middle Powers
An AI strategy for Europe
https://transformative-ai.eu/part-a#a-executive-summary
What should middle powers do?
https://carnegieendowment.org/research/2026/09/artificial-intelligence-ai-development-breakout-germany-india-japan-uk-middle-power
Middle powers will lose AI access sooner than you think
@anton_d_leicht on X
India’s sticking with the US for now
https://www.foreignaffairs.com/india/india-sticking-america-now
Also see this analysis and commentary: @kejimao on X
From the Grapevine
Startups are selling distillation traces
@andriy_mulyar on X