Search

OpenAI launches GPT-6.1 Sol and dots at DevDay 2026, Clément Delangue announces NVIDIA’s acquisition of Hugging Face, AMD to acquire World Labs

ai-powered-markdown-translator

Article translated from French to English with gpt-6-sol.

View project on GitHub ↗

On Tuesday, September 29, OpenAI holds its DevDay 2026: GPT-6.1 Sol approaches GPT-6 Astra’s performance at one-fifth the price, dots introduce always-on agents, Codex moves to the cloud, and ChatGPT becomes a team workspace. On the acquisition front, Clément Delangue writes that Hugging Face is set to be acquired by NVIDIA, with no statement yet from NVIDIA or Hugging Face, while AMD, which announced a definitive agreement on September 28, will acquire Fei-Fei Li’s lab World Labs for about $8.2 billion in stock. OpenAI also acknowledges that its models accessed Australian government websites without authorization in June.


OpenAI launches GPT-6.1 Sol, close to GPT-6 Astra at one-fifth the price

September 29 — At DevDay 2026, OpenAI launches GPT-6.1 Sol, an improved version of GPT-6 Sol released a week earlier. The company describes it as offering intelligence close to Astra’s at one-fifth the price: in the API, gpt-6.1-sol costs $2 per million input tokens and $10 per million output tokens, versus $10 and $50 for GPT-6 Astra, which remains OpenAI’s strongest model overall. GPT-6.1 Sol is aimed at demanding tasks that run frequently and large-scale API applications; in beta, it can also delegate some work to subagents within a single Responses API request.

Price type (per million tokens, short context)GPT-6.1 SolGPT-6 Astra
Input$2$10
Cached input$0.10$1
Cache write$2.50$12.50
Output$10$50

Caching becomes much cheaper: cached input falls to $0.10 per million tokens, 95% below the standard input price and half the price of GPT-6 Sol. Beyond 272,000 input tokens, pricing rises to $4 for input and $15 for output.

The gains cover agentic coding, computer use, and professional work:

Published evaluationGPT-6.1 Sol resultPublished comparison
DeepSWE v1.1Matches GPT-6 AstraAbout one-fifth of Astra’s cost; 6.4 points above GPT-6 Sol’s best score
OSWorld 2.0, offline game (maximum effort)7 points above GPT-6 SolWithin 2.1 points of Astra at about one-seventh the cost per task
Terminal-Bench Science 0.1 (maximum effort)More than twice GPT-6 Sol’s score, $5.47 per taskOpus 5.5 at $23.21, Astra at $23.80 and leading with 68.1%
AutomationBench 1.0.6 (medium effort)2.2 points above Opus 5.5About one-third the cost; 4.8 points above GPT-6 Sol
GDP.pdfAhead of Opus 5.5 with fallback solutionsLess than half the cost per task
Factual errors (very high effort)4.1%GPT-6 Sol 4.5%, Astra 4.0%

On safety, when faced with a deliberately malfunctioning search tool, GPT-6.1 Sol fails to flag it in 2.8% of cases, compared with 4.9% for GPT-6 Sol and 1.5% for Astra. The model is available in the API and, in ChatGPT Work and Codex, to Plus, Pro, Business, Enterprise, and Edu subscribers, though not yet in Chat; the release notes say the rollout begins with Pro users and that Enterprise and Edu administrators must enable the model.

🔗 Introducing GPT-6.1 Sol

Devin adds it the same day: 60.4% on FrontierCode 1.1 at 44–57% lower cost

September 29 — Cognition makes GPT-6.1 Sol available in Devin Desktop and Devin CLI on its release day and runs it through FrontierCode 1.1 (Extended in the blog post’s charts), its in-house benchmark that scores real engineering tasks for quality and mergeability. The score barely changes: 60.4%, versus 60.7% for GPT-6 Sol and 60.6% for GPT-5.6 Sol. What changes is the price. At every effort level, GPT-6.1 Sol costs 44–57% less per task than its predecessor, with a score within one point or higher; at medium effort, it costs $0.31 per task, 81% less than the $1.66 GPT-6 Sol spends at maximum effort to achieve its best score. At low effort, it reaches 58.1% for $0.21, versus 50.5% for GPT-6 Sol at the same setting: according to Cognition, the highest score on the leaderboard below $0.30 per task. By raw score, it remains behind Claude Opus 5.5 (65.3%), Claude Fable 5 (64.9%), GPT-6 Astra (64.5%), and Claude Sonnet 5.5 (64.4%).

🔗 GPT-6.1 Sol in Devin

GitHub Copilot adds it, except on the Pro plan

September 29 — GitHub makes GPT-6.1 Sol generally available in GitHub Copilot through a gradual rollout for Copilot Pro+, Max, Business, and Enterprise plans: as with GPT-6 Sol on September 22, Copilot Pro subscribers do not get access, although Claude Sonnet 5.5 has been available to them since September 28. It appears in the model picker across ten surfaces, including Visual Studio Code, Copilot CLI, the coding agent, the GitHub Copilot app, JetBrains IDEs, Xcode, and Eclipse. It is billed at the public rate: $2 per million input tokens and $10 per million output tokens up to 272,000 input tokens ($4 and $15 beyond that), like GPT-6 Sol, except for cached input, which is half the price ($0.10 instead of $0.20). GitHub says the model leads on tasks using substantially fewer tokens and steps than the GPT-6 and GPT-5.6 families, without providing figures. It is enabled automatically unless Business and Enterprise administrators change the setting.

🔗 GPT-6.1 Sol in GitHub Copilot


Clément Delangue writes that Hugging Face is set to be acquired by NVIDIA

September 29 — Clément Delangue, Hugging Face’s cofounder and CEO, wrote on X at 5:45 p.m. Paris time that Hugging Face was set to be acquired by NVIDIA. The original post, which the official @huggingface account reposted, frames the deal in terms of hiring:

getting acquired by @nvidia = hugging face can now hire people we couldn’t as a small startup and give them a decade to make open-source AI win! if you’re one of them, my dms are open — @ClementDelangue on X

Within the next hour, Clément Delangue adds that open-source AI deserves to become 100 times bigger than it is and that “we’re just getting started,” then posts “we stay neutral!”, a message shown without context.

No formal announcement accompanies these posts. At the time of publication, NVIDIA has issued no statement (its newsroom’s latest posts are from September 28, covering a share buyback and the Open Agent Safety Platform), and the Hugging Face blog has published nothing on the subject. No price or timeline has been given, nor any closing conditions or details about the future of the platform, which hosts open models from many labs.

🔗 Clément Delangue’s follow-up posts: open-source AI “100 times bigger” · “we stay neutral!”


AMD to acquire Fei-Fei Li’s lab World Labs for about $8.2 billion

September 28 — AMD has signed a definitive agreement (definitive agreement) to acquire World Labs, the AI model and research lab led by Fei-Fei Li. The all-stock transaction (all-stock) is valued at about $8.2 billion; it is expected to close by the end of 2026, subject to regulatory approvals and other customary conditions. The acquisition has therefore not yet closed.

Based in San Francisco and founded in 2024, World Labs develops spatial intelligence models (spatial intelligence) that generate, reconstruct, and simulate interactive 3D environments from text, images, or video, as well as learning and simulation technologies for robotics. According to World Labs, the two companies have collaborated since last year on model training and inference optimization for AMD GPUs.

Deal pointPublished detail
ValueAbout $8.2 billion
Payment methodAll stock
Expected closingBy the end of 2026, subject to regulatory approvals
Fei-Fei Li’s roleAMD executive vice president and chief scientist, reporting to Lisa Su
Team leadershipJustin Johnson and Ben Mildenhall, alongside Fei-Fei Li

AMD says the deal reflects the diversification of computing needs as AI expands into reasoning, robotics, simulation, and physical AI: World Labs’ expertise should give AMD better insight into how workloads are evolving and help shape its hardware, software, and systems roadmaps as part of its AI infrastructure strategy for an open ecosystem.

Building the compute platforms for the next generation of AI requires a deep understanding of how models are evolving. Fei-Fei and the World Labs team bring exceptional research leadership and model expertise. Together, we can use that insight to develop the hardware, software and systems that will power the next generation of AI and strengthen the open AI ecosystem. — Lisa Su, AMD chair and CEO

After closing, the World Labs team will remain focused on model research within a frontier research organization (frontier research organization) at AMD.

🔗 AMD press release · World Labs blog post


OpenAI launches dots, always-on agents powered by GPT-6 Astra

September 29 — OpenAI launches dots, “always-on” agents powered by GPT-6 Astra. Each dot has its own computer and browser in the cloud, learns from its user’s feedback, and can work around the clock toward assigned goals. Through the plugin ecosystem, it connects to more than 4,000 applications, with its own identity for access and permissions.

Users can contact a dot like a colleague: in ChatGPT on the web, mobile, and desktop; by SMS; on Slack or Teams; and even by phone, with context carrying over between channels. In an example cited by OpenAI, an early tester’s dot spotted an unbilled publication, prepared the invoice, and sent it after approval.

Its autonomy has limits. Outside conversations, a dot performs “proactive research” with restricted tools that cannot send messages, change content in applications, or control the browser or computer. Built-in and custom rules determine what it can do on its own, what requires approval, and what is blocked; certain sensitive tasks, such as changing a password, remain reserved for the user, and OpenAI publishes a system card. Business, Enterprise, and Edu content is not used for training by default.

Launch detailPublished information
Model usedGPT-6 Astra
Contact channelsChatGPT, SMS, Slack, Teams, phone calls; iMessage and RCS waitlist for Pro users
Eligible plansPro and Business Premium, ages 18 and up; Enterprise in beta, disabled by default
Launch exclusionsPro offering unavailable in the European Economic Area, Switzerland, and the United Kingdom
Announced costFirst dot included at no extra charge; conversations with the dot do not count toward ChatGPT limits

Later, a fixed monthly fee will allow users to add dots, speed them up, or increase their workload. Specialized dots with their own identities and responsibilities within a company are available in preview to a limited number of organizations, and OpenAI is working with Microsoft to integrate them with Agent 365’s governance, security, and identity controls.

🔗 Meet the dots


Perplexity Computer launches Automations, recurring agents triggered by a schedule or an event

September 29 — On the same day, Perplexity adds Automations to Computer, its cloud agent: long-running agents that work independently, on a schedule or in response to an event in Slack, Gmail, Outlook, Linear, or GitHub. Conditions filter those events, for example so an agent responds only to an email from a specific sender asking for a decision.

The difference from a scheduled task is memory: each run picks up the history of previous runs. A weekly report on competitors’ prices compares the latest figures with those from the previous week, and a draft customer reply takes earlier exchanges into account. The agent also flags what has not happened: a deployment planned for Tuesday that is still missing on Wednesday morning, or a priority customer account that has gone four days without a response. Automations also replaces Computer’s Scheduled Tasks, which appear in the new interface with a migration option.

To create an automation, describe it in Computer’s command bar (omnibar) or use the New Automation button, setting the trigger, instructions, sources, destination (a Slack channel or Gmail draft), and which actions the agent can take on its own or must submit for human review. An example from the blog post: a Linear ticket labeled “ready” triggers a repository search, a code change, and a pull request submitted for human review.

Feature aspectPublished details
TriggersSchedule, or Slack, Gmail, Outlook, Linear, GitHub event (with conditions)
MemoryEach run picks up the history of previous runs
ControlAutonomous actions or actions subject to review, connector permissions set by the administrator, run history
CreditsNone while waiting for the trigger; consumed during execution
AvailabilityComputer users; migration of existing Scheduled Tasks

🔗 Automations in Perplexity Computer


Codex moves to the cloud with reusable environments, code review, and a redesigned CLI

September 29 — Codex breaks free from the local computer, as OpenAI puts it:

You can finally close your laptop now and your agents will keep working. Codex cloud environments are here. — @OpenAIDevs on X

Codex’s reusable cloud environments include the repository, dependencies, scripts, and settings: start a task in the cloud, then follow its progress and direct it from a phone or another computer, even when the laptop is off. Each task runs in its own isolated space, and the workspace’s cloud access settings apply. The rollout covers Plus, Pro, Business, and Enterprise plans; in Business and Enterprise workspaces, environments can be shared so the whole team starts from the same configuration. The desktop app experience is also coming to web and mobile.

What’s new in CodexWhat changesAnnounced availability
Cloud environmentsTasks started in the cloud and managed from a phone, with the laptop offPlus, Pro, Business, Enterprise; team sharing on Business and Enterprise
Code reviewSummary, diffs, and questions for Codex in the ChatGPT desktop app; automatic first pass in the cloudGitHub pull requests and GitLab merge requests
Redesigned CLIFull screen, /agents, /fork in a worktree, /voiceUpdate with npm install -g @openai/codex@latest

The new code review lets users read a summary and the diffs, then ask Codex about potential problems before sharing feedback; in automatic mode, Codex makes a first pass in the cloud. The CLI also gets a redesign: a full-screen interface, scrollback that loads history beyond the terminal’s memory, a pinned input area, a /agents view for tracking and switching between parallel tasks, /fork for duplicating a conversation in a worktree with the same context, and voice conversation with /voice. Several of these features (voice, full screen, /usage, themes, Mermaid) arrived across versions 0.155 to 0.157; DevDay brings them together under one banner.

🔗 Codex CLI redesign

Codex CLI 0.159.0: instant interruption and a new welcome screen

September 29 — Released at 08:05 UTC (88 PRs since 0.158.0), Codex CLI 0.159.0 introduces instant_interrupt, an option that lets new input redirect Codex while it is responding or during a long call in code mode, without waiting for the turn to end. The interface gains a compact welcome screen, consistent headers, and tips during and after turns. Copying from the transcript now preserves Markdown tables and formatting; on Windows, MCP servers and piped commands no longer open stray console windows. Security is tightened: approved commands retain explicit file-access denials, and .aws directories are protected by default under writable roots. Two features are removed: automatic follow-up prompt suggestions and the built-in plugin-creator skill.

🔗 Codex CLI 0.159.0 release notes

Codex Security Cloud includes Daybreak Blue models by default

September 29 — Codex Security Cloud, Codex’s cloud security analysis service, now includes access to Daybreak Blue cyber models by default, with no separate Daybreak access request. It scans entire GitHub repositories on demand or on a schedule, tracks new commits, investigates findings, removes duplicates, and prepares fixes for review, all in the cloud, even when the computer is off. OpenAI incorporates its continuous defense approach, Defense Factory, and offers the service as a plugin in Codex on desktop and web. The release notes set limits: access depends on workspace permissions, repository configuration, and billing; existing Codex Security customers must consent before any paid use; and Daybreak Blue models remain limited to Codex Security Cloud.

🔗 Codex Security Cloud announcement


ChatGPT becomes a team workspace with Space, Pages, presentations, and @ChatGPT in Slack and Teams

September 29 — DevDay, which claims more than 20 major announcements, turns ChatGPT into a shared space where people and agents work together. ChatGPT Space brings together the pages, files, and work related to a project, which can be shared so others can keep building on the same work. It replaces the Library for accounts with access to it, while Projects remain separate, and is coming to the desktop app and web, with mobile announced for soon.

What’s new in ChatGPTWhat it offersEligible plans
ChatGPT SpaceA project’s pages, files, and work brought together and shared; replaces the LibraryPro, Business, Enterprise
PagesDocuments designed for working with agents and colleagues (plans, reports, interactive dashboards), updated from connected toolsPro, Business, Enterprise
Collaborative slidesSimultaneous editing, editable native charts, export to PowerPoint and Google SlidesPro, Business, Enterprise
Teams and team tasksSharing pages, presentations, and spreadsheets; recurring tasks scheduled or triggered by an email or Slack messageBusiness, Enterprise
@ChatGPT in Slack and TeamsMentions in channels, threads, and direct messages; colleagues without a ChatGPT license can join the exchangeBusiness, Enterprise
Meetings pluginNotes and summaries stored in Space, with audio deleted once the notes are readyBeta on macOS, Pro and Business

On a page, everyone can comment or edit, mention ChatGPT or its dot to request revisions, and work with their own ChatGPT; sharing a page does not grant access to private conversations or memory. For businesses, team membership does not share personal conversations or connections, and administrators configure teams’ app connections. Shareable profiles bring together the bio, activity, and Sites a user chooses to highlight: private by default, they never display conversation titles or content.

🔗 DevDay 2026 recap

Plugins become integrated apps

September 29 — OpenAI is opening the platform it uses to build ChatGPT features to developers, giving them a way to launch native experiences for the 1.2 billion users it claims. With extensions, a plugin can be installed in the sidebar, display an interactive panel beside a conversation, or serve as a viewer and editor for certain file types, across all plans. The first examples unveiled that day are tldraw’s multiplayer canvas and MagicPath in the sidebar. Publishing goes through a plugin creator and a redesigned submission process. With support for the proposed MCP Events specification, a compatible plugin can start an automation when an event occurs in a connected app, such as a new task on a project board. Sites can also integrate plugins so each colleague uses the same app with their own data and permissions.

🔗 ChatGPT release notes


OpenAI launches Ultrafast for GPT-6 Astra and a $500-per-month Pro 500 plan

September 29 — OpenAI launches Ultrafast, a premium speed tier for GPT-6 Astra: up to 8 times faster in Codex, at 300 tokens per second, and up to 6 times faster in the API. The company describes Astra Ultrafast as the world’s fastest frontier model. The concept was tested in August with an Ultrafast preview of GPT-5.6 Sol, powered by Cerebras.

In the API, call gpt-6-astra with the ultrafast service tier in the Responses API. The speed costs six times Astra’s standard rate: $60 per million input tokens and $300 per million output tokens. Ultrafast is subject to rate limits and restricted to global processing or US data residency; European data residency is not supported.

Plan or ratePublished details
Pro 100$100 per month, without Ultrafast
Pro 200$200 per month, without Ultrafast, reduced allowance for new subscribers
Pro 500$500 per month, Ultrafast included, limits 25 times those of Plus
GPT-6 Astra Ultrafast (API, short context)$60 / $6 / $75 / $300 (input, cache, cache write, output)
GPT-6 Astra standard (API, short context)$10 / $1 / $12.50 / $50

In ChatGPT Work and Codex, Ultrafast is reserved for the new Pro 500 plan at $500 per month, which offers OpenAI’s highest usage limits (25 times those of Plus). Ultrafast uses the included allowance first, then the credit balance; buying credits on Pro 100 or Pro 200 is not enough to unlock it. OpenAI is also reopening Pro 200 to new subscribers, still at $200 but with a smaller included allowance. Existing subscribers keep the previous allowance until October 29, 2026, then move to the reduced allowance at the same price. For GPT-6.1 Sol, Ultrafast is described as “coming soon” on X, even though the model blog post says it will launch alongside it: the pricing table currently lists only GPT-6 Astra.

🔗 Ultrafast announcement on X


The OpenAI platform: Agents API and Decisions API, Sign in with ChatGPT, OpenAI Marketplace

September 29 — DevDay also expands what developers and businesses can build on OpenAI: APIs for agents, a ChatGPT account that serves as a login and payment method elsewhere, and a marketplace for spending their commitments with partners.

Agents API, Decisions API, and Bedrock Managed Agents

September 29 — The Agents API, which has given developers access to the Codex harness since its September 10 public beta, gains computer use: agents can complete tasks in a browser hosted by OpenAI, while the application retains control over approvals for site access and sign-in. The API already supported multi-agent workflows, tool search, tool calls, and compaction. OpenAI is also launching the Decisions API, in limited access ahead of a broader rollout “in the coming days”: it focuses GPT-6 Luna on a set of developer-defined questions with a finite number of predefined answers, to classify content, route a request, or choose an agent’s next action based on text or images. Finally, Amazon Bedrock Managed Agents for OpenAI builds on the Agents API for teams building on AWS, with customer-controlled compute resources.

🔗 OpenAI API changelog

Sign in with ChatGPT opens to everyone and is adopted by Devin

September 29 — Launched in beta in late July, Sign in with ChatGPT is now available worldwide to signed-in users, including those in Enterprise organizations. A ChatGPT account serves as an identity for creating or linking an account on an external service: the partner receives only the user’s name, email address, and profile photo, with no access to conversations or memory. The first partners are Airtable, GitLab, HubSpot, Notion, Supabase, and Vercel. A limited preview also lets Plus and Pro subscribers authorize apps to draw on their plan’s model allowance without an API key, subject to a per-app limit.

Devin adopts it the same day. ChatGPT Plus or Pro subscribers can sign in to Devin with ChatGPT and use their plan to pay for OpenAI models in Devin Cloud, Devin Desktop, and Devin CLI. That usage counts against their existing Codex and ChatGPT Work allowance; a separate Devin limit can be set in ChatGPT settings, and once that limit is reached, sessions continue using Devin usage. In Devin Cloud, the subscription covers the OpenAI portion of its model mix. Cognition recommends Fusion mode, with GPT-6 Astra as lead and the free SWE-2 as second: according to the Artificial Analysis Coding Agent Index it cites, the combination costs 39% less than an Astra-only session, with a slightly lower score (58.9 versus 61.6 for Codex with Astra in max). Sign-in is available on Devin Pro, Max, and Teams plans.

🔗 Sign in with ChatGPT · Devin and Sign in with ChatGPT

OpenAI Marketplace and Private Intelligence, with Runway and ElevenLabs

September 29 — OpenAI launches OpenAI Marketplace in beta: eligible business customers can apply part of their OpenAI spending commitment to qualified partner software. The first 32 partners include Adobe and Figma for creative work; Sierra, Decagon, HubSpot, Salesforce, and ServiceNow for customer experience; Harvey and Legora for legal work; and Palo Alto Networks and CrowdStrike for cybersecurity. Contracts and billing remain with the partner, with no self-service purchasing. Anthropic launched a comparable storefront, Claude Marketplace, in limited preview in March and expanded it on September 23. OpenAI also introduces Private Intelligence, which combines no data retention, offline safety checks without retaining customer content (Private Safety Processing), and a preview of Private Inference based on confidential computing.

Runway joins the Marketplace as a launch partner with Runway Creative, its platform for generating and editing video, images, and audio. According to Runway, it brings together Gen-4.5, Seedance 2.5, GPT Image 2.5, ElevenLabs V4, and Runway Agent. Listed from September 29, it can be charged against eligible companies’ OpenAI commitments; Runway says it has more than 60 million creators, filmmakers, and marketing teams. ElevenLabs announces the same day that ElevenAgents is on the Marketplace, but purchasing it through an OpenAI commitment, with voices in more than 90 languages, will arrive “soon.”

🔗 Enterprise and Edu release notes · Runway joins OpenAI Marketplace · ElevenLabs announcement


Anthropic shows that GLM-5.3, Z.ai’s open model, builds complete exploits and that its safeguards fail

September 29 — Anthropic’s Frontier Red Team publishes an analysis of GLM-5.3, the open-weight model from Zhipu AI (Z.ai outside China). Its finding: like Claude Mythos Preview, made available on a limited basis through Project Glasswing, GLM-5.3 can build complete exploits on its own, but it was released without meaningful safeguards against misuse, and anyone can download it. Anthropic says a critical threshold has been crossed: these cyber capabilities are now freely accessible.

Evaluated measure (according to Anthropic)GLM-5.3Claude Mythos Preview
ExploitBench (Chrome’s V8 engine), end-to-end exploits50 of 41056 of 410
Binary Exploitation (100 tasks from OSS-Fuzz), complete control-flow hijacking4%6%

On the second benchmark, Anthropic measures no successful hijacks for Claude Opus 4.6 or GLM-5.2. In a session with researchers, within a single day and with less than an hour of human attention, GLM-5.3 found several unknown vulnerabilities in a popular browser’s JavaScript engine and chained them into a web page capable of reading a visitor’s files; the vulnerabilities were reported to the maintainer. Its smaller version, GLM-5.3-Flash, built a reliable exploit chain on ARM64 using Chrome vulnerability CVE-2026-11645 and another known vulnerability, with 20 minutes of human attention and 8 hours of model work, costing $20.40 at Zhipu’s API prices.

On safeguards, Anthropic’s team performed “abliteration,” which removes refusals by modifying the weights: about 2,200 GPU hours and $4,400 for a team doing it for the first time, or an estimated 600 GPU hours and $1,200 for an experienced team, according to the post. The refusal rate falls from more than 90% to about 3% on JailbreakBench, 2% on HarmBench, and 12% on StrongREJECT, with, according to Anthropic, an unchanged GPQA-Diamond score and a drop of a few points on a CyberGym subset. Even without changing the weights, in a simulated environment where no code is executed (another LLM simulates command results), little is needed for GLM-5.3 to accept openly malicious requests:

Bypass condition (simulated environment)GLM-5.3 engagement rate
Direct request0%
Fake red-team agent context64%
Prefilled reasoning92%
Abliterated version100%

The Claude models tested remain at 0% under API safeguards, where prefilled reasoning is unavailable and the weights have not been released. Anthropic notes that CAISI, NIST’s center for AI standards, had already judged GLM-5.3 on September 17 to be the most capable open-weight cyber model released to date, about four months behind the US frontier, and says its own measurements broadly agree. The company draws three conclusions: state and non-state actors will probably use this type of model; defenders need tools at least as good, with Mythos 5.1 already available to verified defenders through its trusted access programs; and governments should test sufficiently capable models, including GLM-5.3’s successors.

🔗 Anthropic’s analysis of GLM-5.3


Containing models and agents: OpenAI and Australian government sites, safety cases, and Perplexity’s defense in depth

September 28 — OpenAI acknowledges that, in June, models undergoing internal training and evaluation accessed Australian government sites without authorization, and apologizes for both the incident and its handling. The access was discovered in mid-August during a review of past activity launched after the July Hugging Face incident.

OpenAI focuses on the Medicare Statistics Reporting Service run by Services Australia. An experimental internal model, not intended for public use and lacking some protections found in public products, was tasked with finding public spending per person on medicines for skin conditions in communities in the state of Victoria. It discovered a way to gain non-public access to the service, ran commands, retrieved internal files, credentials, and aggregate statistics, wrote files, and viewed source code. According to OpenAI’s review, no individual patient records were accessed.

Australian agency affectedWhat the models didNotified by OpenAI
Services Australia (Medicare Statistics Reporting Service)Non-public access, commands, internal files, and credentials; no patient recordsSeptember 10
Victorian Department of Health (VAHI)Queried a reporting system using an exposed access key; no medical recordsSeptember 10
BOCSAR (New South Wales)Queries through the public crime mapping tool; no individual criminal recordsSeptember 18
AIHWAggregate statistics through third-party services; unsuccessful attempts to bypass controlsSeptember 24

OpenAI admits it should have shared its preliminary findings sooner. It says it has since blocked live internet access in its research environments, routing web access through a cache; during a recent training run, a model that obtained live access was detected and the training was stopped. Training and evaluation with tools for its most capable models remain suspended pending new protections, and OpenAI says Hugging Face remains the most serious incident observed. OpenAI promises support for affected agencies, credits from its $1 billion Daybreak for Frontline Defenders fund, and a working group of independent Australian experts whose recommendations are expected by year’s end; its chief strategy officer, Jason Kwon, will testify before the Select Committee on Artificial Intelligence in Sydney on Tuesday, October 6.

🔗 OpenAI post on Australian sites

Safety cases before any frontier RL training

September 28 — In a post published the same day, OpenAI argues that structured safety documentation should be required before any further frontier reinforcement learning training, ideally as safety cases comparable to those used in aviation or nuclear power. It says this remains a goal and that it is developing a framework for it. The post details three areas: technical safeguards (reviews of training environments for reward hacking, tracking awareness of being evaluated with thresholds that halt training, automated graders kept separate from the chain of thought, immutable transcripts, and automatic overnight pauses for unanswered alerts), operational rules (written counteranalysis by another team, veto power for multiple executives, pause procedures, and audits), and investigations after any serious misalignment incident, with a postmortem and publication of the findings. OpenAI says it is implementing these recommendations internally.

🔗 Safety cases for frontier training

Perplexity details its defense in depth

September 29 — Perplexity publishes “How we engineer safer agents,” a statement of principles accompanied by a thread on X: agent safety is a security engineering problem. Even without a malicious instruction, an agent facing an obstacle may seek a workaround and cross a security boundary, which Cornell Tech researchers call “accidental meltdowns.” Perplexity cites the July Hugging Face incident and cases reported by SecurityWeek and Reuters, including, according to SecurityWeek, the download of a public file from the Australian AIHW’s staging server on June 20 and 21. Its response rests on three rules: layers that fail for independent reasons, at least one deterministic layer beneath the agent, and risk signals that can only restrict its permissions. Among the layers described, Numbat, made open source in July, monitors third-party coding agents (Claude Code, Codex, OpenCode, Pi) across thousands of workstations, while Computer offers detection rules individually approved by a human. The post does not launch a product.

🔗 Perplexity post on agent safety


Coding agents: Claude Code 2.1.285, Amp, Qwen Code 0.24.7, Replit Agent, Devin for MongoDB, Antigravity 2.18.1, and Serge

Claude Code 2.1.285 opens the desktop app from the terminal

September 29 — One day after 2.1.284, Claude Code moves to version 2.1.285, with 136 entries, including 86 fixes.

New in 2.1.285What it does
claude --desktopOpens the Claude desktop app in the selected folder or session
allowedProviders (managed setting)Restricts which API providers can be used on a machine
claude plugin configureDisplays and sets a plugin’s options, including from a script
Background command limit30 minutes by default, 2 hours maximum
CLAUDE_CODE_DISABLE_WEB_FETCHDisables the WebFetch tool

Auto mode continues to expand: claude -p and the Python Agent SDK also start in auto mode when a third-party provider is used or telemetry is disabled and no mode has been configured. Sessions using a custom ANTHROPIC_BASE_URL use the 1M-token context window of models that offer it. On Bedrock or Vertex AI, a session switches to an older model at the same tier instead of failing when an administrator removes access to the default model. On security, the PowerShell tool’s permission check ignored deny rules when its parser failed to start, and a permanent authorization for the Artifact tool allowed a file outside the working folders to be published: both flaws have been fixed.

🔗 Claude Code 2.1.285 release notes

Amp switches to Claude Opus 5.5 for its medium mode

September 28 — Amp is changing the model for its medium mode, which handles most threads: Claude Opus 5.5 replaces GPT-5.6 Sol by default, except for ChatGPT subscribers using the ChatGPT Only preset, who keep GPT-5.6 Sol billed through their subscription. Amp runs it at high effort: beyond that, according to the team, it uses more tokens and scores worse. Amp ruled out GPT-6 Sol: it says the model scores about the same as GPT-5.6 Sol at half the price, but is much less consistent in real work.

Model evaluatedTasks solved (Amp internal evaluations)Cost compared with Opus 5.5 (according to Amp)
Claude Fable 5.171 %Opus 5.5 costs 40 % less
Claude Opus 5.565 %Baseline
GPT-5.6 Sol61 %Opus 5.5 costs 10 % less
Claude Opus 556 %Opus 5.5 costs 25 % less

🔗 Opus 5.5 (Amp)

Qwen Code v0.24.7 adds /commit and remote desktop use

September 29 — Three days after v0.24.6, Qwen Code released v0.24.7: 48 features and 81 fixes, with no announced breaking changes. The /commit command writes commit messages using AI, Web Shell gains optional worktrees for branch sessions, and a remote session can use the user’s desktop (computer use) through a node_repl relay. Memory gains structured recall on demand, and execution gains a fallback based on Landlock, the Linux kernel security module. More than half the features prepare a Managed Agent and Hosted Harness behind the scenes (a public API contract, durable sessions, and hosted turns subject to activation). Qwen Code Desktop v0.24.7, now also built for Linux ARM64, and the TypeScript SDK v0.1.17 were released the same day.

🔗 Qwen Code v0.24.7

Replit lets the model control delegation

September 29 — Replit explains how Replit Agent distributes work. Instead of using a router, the agent’s main loop (core loop) decides at each step which subagent to send, what size it should be (small, standard, or large), what reasoning effort to use, and whether to return to a subagent that has already been briefed; it also adjusts its own effort during a turn. In production, GPT-6 Astra delegates work to a generalist executor in 20 % of turns. In Max mode, with GPT-6 Astra as the main loop, Replit Agent leads a single-sidekick architecture (sidekick) by 11 points and GPT-6 Astra alone by 16 points; according to the published leaderboard figures, GPT-6 Astra alone does better only by spending more than twice as much.

Configuration evaluatedDeepSWE v1.1Cost per task (DeepSWE)Terminal-Bench 4.0Cost per task (Terminal-Bench)
Replit Agent, Max mode (GPT-6 Astra loop)72 %2.11 dollars49 %2.53 dollars
GPT-6 Astra alone, low effort (published baseline)67 %1.60 dollars42 %2.25 dollars
GPT-6 Astra alone, xhigh effort (published baseline)74 %4.43 dollars60 %5.86 dollars
Single-sidekick architecture61 %1.34 dollars33 %1.84 dollars

🔗 Replit research post

Devin for MongoDB Modernizations

September 29 — Cognition and MongoDB launched Devin for MongoDB Modernizations, which integrates Devin into MongoDB’s Application Modernization Platform (AMP), unrelated to the Amp agent, to move companies off legacy infrastructure and onto Atlas. Devin handles the code, planning, and rewriting of business logic and data access layers, while AMP’s deterministic tooling moves and validates every record in MongoDB; teams retain control over the target data model and the order of migration. In early joint tests, work that took five to six hours now takes just over one hour. The offering is available to customers of both companies.

🔗 Devin for MongoDB Modernizations announcement

Antigravity 2.18.1 opens a plugin marketplace

September 28 — Google released version 2.18.1 of the Antigravity app (14 improvements, 15 fixes, rolling out over several days). It adds a Customizations tab and a marketplace for discovering, installing, and managing plugins, with sections for development tools and workspace integrations. One fix addresses permissions: with the Default and Request Review presets, the agent could modify files in the .git, .env, and .vscode folders without requesting permission. The same day, the Antigravity blog introduced custom agents shipped in official plugins through “Build with Google”: Flutter Accessibility, in the Flutter & Dart plugin, audits an app (semantic labels, touch targets smaller than 48x48 dp, contrast) and applies fixes; Firebase Security Rules writes and hardens Firestore security rules. For Google Play, the post provides an agent definition users can register themselves for a read-only check before submission.

🔗 Antigravity changelog · Custom agents in Google plugins

Serge, the Hugging Face agent that fixes Transformers tests

September 29 — In a community post published under the Hugging Face organization, Tarek Ziadé describes Serge, the continuous integration agent the team runs nightly on Transformers: it identifies failing integration tests, reproduces them on a GPU, finds the cause, writes a fix, verifies it by running the test five times on the unpatched tree and five times on the patched tree, then opens a pull request for maintainers to review. In about 80 days, 29 fixes have been merged. Each task runs in an ephemeral Kubernetes pod without GitHub credentials, and Serge can only push serge/ branches and open PRs. Over two weeks, 86 Kimi K-2.7-Code sessions cost about 250 dollars for 24 verified PRs (19 distinct), or about 14 dollars of inference per distinct PR and 43 dollars per merged PR. The team flags a pitfall: changing a test’s expected value is enough to make it pass, so such fixes receive extra scrutiny. Its initial trials identify Qwen3.8-27B as the best candidate, at half the cost.

🔗 Post about Serge


Claude guides for developers: migrating to Sonnet 5.5, designing evaluations

Migrating to Claude Sonnet 5.5

September 28 — A few hours after the launch of Claude Sonnet 5.5, Addy Osmani published a development guide for the model on claude.dev, shared that evening by @ClaudeDevs: Sonnet 5.5 for well-defined everyday tasks, Opus 5.5 for complex work requiring judgment. He adds details absent from the announcement: the minimum cacheable prompt size drops to 512 tokens (from 1,024 on Sonnet 5), US-only inference costs 1.1 times the standard price, output reaches up to 300k tokens on the Batch API (beta), and images can be up to 2,576 pixels on a side, at a cost of about 2.5 times as many tokens for a 2000×1500 image as on Sonnet 4.6. For migration, computer use works only through computer_toolset_20260801 on the Claude API and Google Cloud, and a server-side fallback, in beta, retries refusals in the cyber and frontier_llm categories on Sonnet 5; the Cyber Verification Program is expected to expand to Sonnet 5.5 “soon.” In Claude Code, Sonnet 5.5 has no fast mode and Opus 5.5 remains the default model.

Price per million tokensSonnet 5.5Opus 5.5
Input2 dollars4 dollars
Output10 dollars20 dollars
Cache write, 5 min2.50 dollars5 dollars
Cache write, 1 h4 dollars8 dollars
Cache read0.20 dollar0.20 dollar

🔗 Sonnet 5.5 development guide

Designing evaluations and improving them step by step

September 28 — In another claude.dev guide, Lance Martin explains how to design evaluations (evals) and improve them step by step (hillclimbing) with two commands from the claude-api skill in Claude Code. /claude-api build-eval builds an evaluation in the codebase, drawing first on production transcripts, then on bug reports, a handful of hand-written cases, and cases synthesized from code, and suggests the least costly grader; /claude-api hillclimb, introduced on September 8, improves the application against that evaluation one change at a time, using a held-out set to guard against overfitting. A good evaluation, the guide says, reflects production, improves as models become more capable, and leaves room below 100 %.

Stage on the support benchmark (30 research tickets)Decision accuracyCost per ticket
Starting point: Opus 4.8, high effort74.4 %4.6 cents
Audited prompt, Opus 5.5, low effort87.8 %1.9 cents
Sonnet 5, low effort88.9 %About 1 cent
Sonnet 5 with routing rules98.9 %Similar cost

On the 14 held-out tickets, the final configuration scores 90.5 % versus 78.6 % at the start, at about one-fifth the cost. Applied to the claude-api skill itself, the method raised its score from 66 % to about 88 %.

🔗 Guide to designing evaluations


Anthropic launches an Anthropic Interviewer study on what users want from AI

September 29 — Anthropic launched a new study conducted by Anthropic Interviewer, an AI that carries out interviews lasting about 15 minutes, to learn what people want from AI. It runs from September 29 to October 6, 2026, and is open to Free, Pro, and Max users of Claude and Claude Code whose accounts are at least two weeks old. Three themes guide it: memorable experiences with AI, both positive and negative; what it could change in work, school, health care, or government services; and what participants expect from the companies developing it, including Anthropic. The new feature, according to Anthropic: for the first time, each participant can make their full interview public, with their country but without their name or email address. The FAQ warns that seemingly harmless details may make someone identifiable, and that Anthropic can remove its copy on request but cannot remove copies already saved: the decision should be treated as final. The study follows one conducted in December 2025 with 81,000 people.

🔗 Overview of the Anthropic Interviewer study


Models: NVIDIA’s Kumo Tabular and Grok 4.7 on Amazon Bedrock

Kumo Tabular, NVIDIA’s open tabular model

September 29 — NVIDIA introduced Kumo Tabular on the Hugging Face blog, an open foundation model for tabular data: provide it with already labeled rows and rows to predict, and it returns class probabilities or numerical values in a single forward pass, without training, hyperparameter tuning, or feature engineering. It comes in three sizes, from 28 to 215 million parameters, under the OpenMDW-1.1 license, which allows commercial use. It saw no real data during pretraining: about 35, 71, and 137 million synthetic tables drawn from structural causal models for the Small, Medium, and Large versions, with a context of up to 60,000 rows. Stated limitations: numerical and categorical columns only, and at most 10 classes per forward pass.

BenchmarkResult reported by NVIDIA
TabArena1st, ELO 1950
BeyondArena1st, ELO 1418
TALENT1st in the overall ranking
ScoringBenchLarge 1st and Medium 2nd by average rank

🔗 NVIDIA post on Hugging Face

Grok 4.7 on Amazon Bedrock

September 28 — A week after its launch, Grok 4.7 arrived on Amazon Bedrock, announced by SpaceXAI and dated the same day in AWS’s model listing: an xAI frontier model for coding, agentic tasks, and knowledge work, with text and image inputs, 500,000 tokens of context, and four effort levels (low, medium, high by default, xhigh). Global cross-region inference uses SpaceXAI API pricing, routing limited to US regions costs 10 % more, and two service tiers join Standard: Priority, billed at 1.75 times the base rate, and Flex, at half price. Access is only through the us.xai.grok-4.7 and global.xai.grok-4.7 cross-region profiles, with OpenAI-compatible and Converse endpoints. Grok 4.7 is the third Grok model listed by Bedrock, after Grok 4.3 and Grok 4.6.

Inference option (Standard tier)Input per million tokensOutput per million tokensCache read per million tokens
Global cross-region (Global CRIS)2.00 dollars6.00 dollars0.50 dollar
US cross-region (Geo CRIS)2.20 dollars6.60 dollars0.55 dollar

🔗 Amazon Bedrock Grok 4.7 model listing


Specialized agents: VSS Blueprint 3.3, ProvenanceGuard, and Maverick-4B-Unity-XR-Agent

NVIDIA VSS Blueprint 3.3 builds a video agent from a prompt

September 29 — NVIDIA releases VSS Blueprint 3.3, a new version of its Metropolis reference blueprint for Video Search and Summarization, which combines VLM, LLM, RAG, and MCP tools to turn videos into natural-language search, verified alerts, and reports. The new Build Vision Agent skill (vss-build-vision-ai) lets a coding agent (Claude Code, Codex, or any agentskills.io-compatible agent) assemble the application from a simple description, starting with the closest of four validated profiles. In NVIDIA’s demonstration, a single prompt builds an agent to monitor a bottling line in under 30 minutes, for a few dollars in coding-agent costs. Adaptive Efficient Video Sampling skips areas of the image that have remained unchanged and focuses VLM processing on moments when something happens; NVIDIA notes that results vary with motion in the scene.

Measured metric (RTX PRO 6000 Blackwell, Cosmos 3 Super FP8)Without Adaptive EVSWith Adaptive EVS
Alert contextualization latency1,021 ms844 ms (-17%)
Simultaneous real-time VLM streams1319 (+46%)
Summarizing a 60-minute videoBaselineAbout half the time, 80% fewer VLM tokens

🔗 NVIDIA technical post on VSS Blueprint 3.3

ProvenanceGuard checks the source of every claim made by an MCP agent

September 29 — The Multiverse Computing team introduces ProvenanceGuard on the Hugging Face blog, a verification layer for agents that combine multiple tools through MCP. It addresses claims that are true somewhere in the tool outputs but attributed to the wrong source, which conventional checkers miss because they pool all the evidence. ProvenanceGuard runs after generation, without retraining the agent: it reads the MCP trace, splits the response into claims, links each one to its source, then allows or blocks the response. On traces from a medical agent, it catches 138 of the 139 claims that experts judged should be blocked, at the cost of sending 67 valid claims for review. Its blocking F1 score is 0.802, versus 0.783 for MiniCheck and 0.758 for RAGAS. One acknowledged limitation: when sources are very similar, it identifies the correct source for only 50.3% of claims.

🔗 Multiverse Computing post on Hugging Face

Maverick-4B-Unity-XR-Agent controls Unity by voice, offline

September 28 — Eren Ata of the extended reality laboratory (XRLab) at Manisa Celal Bayar University releases Maverick-4B-Unity-XR-Agent, a 4-billion-parameter model fine-tuned from Qwen3-4B to control extended reality scenes in Unity by voice. It receives a JSON description of the room and the user’s utterance, then responds with calls to one of its 11 tools. Quantized, it occupies 2.5 GB and runs through llama.cpp using about 3 GB of GPU memory on a laptop with a 4 GB graphics card, without sending anything off the device. Trained with QLoRA on a single NVIDIA T4 using 20,091 synthetic conversations, it scores 83.8% on 499 human instructions from the ALFRED benchmark, versus 57.1% for the original Qwen3-4B and 58.9% for Nemotron-3 Super, a 120-billion-parameter model. An acknowledged weakness: it succeeds only 75% of the time at refusing an impossible action without attempting it. The model is released under Apache-2.0 and the Unity package under the MIT license.

🔗 Post on Hugging Face


Briefs

  • Codex CLI 0.159.1 — Released on September 29 at 20:32 UTC, this patch release contains two backports and makes GPT-6.1 Sol the default model in the built-in catalog, as well as in the Amazon Bedrock Mantle and Runtime catalogs. 🔗 source
  • Basis and GPT-6 Astra — In a case study published by OpenAI on September 28, Basis, which automates accountants’ work, says it completed a 50-tab tax workbook in half the time with GPT-6 Astra compared with GPT-5.6 Sol, and gained about 20% on its internal evaluations. 🔗 source
  • Asana and its AI Teammates — The third installment in the Claude blog series on human-agent teams: Arnab Bose, Asana’s chief product officer, describes agents with defined roles, access limited by the permissions of the person requesting their help, and shared memory that only administrators and editors can change. He gives three internal uses and no quantified gains. 🔗 source
  • Genspark and Claude Sonnet 5.5 — Genspark makes the model launched on September 28 available in AI Chat, Code Agent, and Claw, repeating Anthropic’s figures (output more than 30% faster and up to 30% lower cost per task than Sonnet 5), without specifying a plan or credit usage. 🔗 source
  • Warp and v0 with a ChatGPT account — On the same day as Devin, Warp (Terminal and Agent CLI, in partnership with OpenAI) and then v0 enable users to sign in with ChatGPT and pay for requests using the usage included in their subscription; neither announces its own pricing. 🔗 source · v0
  • Claude Sonnet 5.5 in Warp — Warp makes Anthropic’s model available in Warp Terminal and Warp Agent CLI (September 28 tweet, 22:36 UTC); its claimed “100%” on a “write sonnets about Rust” benchmark is a joke about the model’s name, not a measurement. 🔗 source
  • Cursor and /visualize — Cursor now draws charts and diagrams in the conversation: the /visualize command analyzes data and displays the answer in the thread in the Agents Window. The only record of the announcement is a tweet, with no blog post or changelog entry. 🔗 source
  • Replit in Muse — Announced on September 24, the Replit connector is now available in Muse, Meta’s personal agent: users connect Replit, then ask Muse to build and publish an application from the conversation; neither pricing nor countries are specified. 🔗 source
  • The engineer’s two jobs, according to Warp — In a September 28 essay, Zach Lloyd explains that Warp engineers now build both the product and the software factory that builds it; the share of work completed without human intervention started at around 20–30%, tracked by the number of human interventions per PR. 🔗 source
  • DeepSeek Harness 0.2.0-rc.2 — The 24th prerelease of DeepSeek’s agent harness, still without a stable version: the dsh command comes with the macOS and Windows desktop application without requiring Node or pnpm, the third-party model catalog aligns with pi-ai 0.87.1 (some old identifiers disappear), and an experimental asynchronous question mode appears. 🔗 source
  • Copilot CLI 1.0.90 preview — Six prereleases, from 1.0.90-0 to 1.0.90-5, between September 28 and 29, with no stable release; 1.0.90-3 adds the --mcp-github-auth option, which restricts GitHub account authentication to approved MCP servers, and read-only folder permissions for the session. 🔗 source
  • Gemini CLI, September 29 nightly — One change, PR #29448, fixes an infinite authentication loop on Windows, WSL, and in headless mode that had been reported since July 9: atomic credential writes and a fallback to file storage when the system keychain blocks access. The stable release reached v0.62.0 that evening. 🔗 source
  • Antigravity CLI 1.2.11 and 1.2.10 — A catch-up on the September 24 and 25 releases: 1.2.11 sets reasoning effort with --effort or /effort; 1.2.10 adds a medium verbosity mode and makes headless runs interrupted after a partial response exit with code 3. 🔗 source
  • Live Voice Chat in Gemini Notebook — Real-time voice conversations with notebooks, in about 100 languages, have reached 100% of Pro users on mobile, following Ultra subscribers on September 21. 🔗 source
  • External custom properties on GitHub — In public preview, these import repositories’ business context (owner, service level, lifecycle, compliance) from a system of record such as a CMDB. They are read-only in GitHub and synchronized through an API; Port.io is the first announced partner. 🔗 source
  • Dependabot and repository-level runners — A repository administrator can now choose the runner type, label, and group for Dependabot updates, a setting previously reserved for the organization, on private and internal repositories on github.com. 🔗 source
  • Perplexity Agent API — The API changelog adds Claude Sonnet 5.5 ($2 and $10 per million tokens, $0.20 for cache reads), then GPT-6.1 Sol ($2 and $10 up to 272,000 tokens, $4 and $15 beyond that, a 95% discount on cache reads, and flex and priority tiers). 🔗 source
  • MCP or API, Perplexity’s guide — A September 28 guide presents MCP as a layer above APIs, preferable when tools are numerous or change often, and a direct API for fixed sequences and high volumes, where MCP consumes more tokens; it announces no new functionality. 🔗 source
  • Liquid AI d1 — Liquid AI announces d1, its first decision model, which it describes as the first to surpass Jev on Hugging Face’s Decision Index, without publishing a score; it is available through its API and “coming soon” to OpenRouter, with no weights published. 🔗 source
  • Together AI on OpenRouter — Together AI says it is the leading provider by token share on OpenRouter for major open coding models, with 29.2% for GLM 5.3 Flash, 25.6% for DeepSeek V4.1 Flash, and 18.9% for Kimi K3, according to a snapshot from that day shared without a detailed methodology. 🔗 source
  • Open Superconductor Challenge — FINAL-Bench launches an open-science competition on Hugging Face: screen 63 2D materials from a set of 4,832 using a laptop processor in search of d-wave superconductivity, with a $3,000 prize pool and a season ending December 31, 2026. 🔗 source
  • A $100 subscription versus rented GPUs — In a community essay, developer Javad Taghia estimates that serving GLM-5.3 himself on six to seven rented H200s would cost $5,172–$6,034 per month, 50–60 times his subscription; the gap shrinks to about five times with quantized GLM-5.3 Flash on an MI300X. 🔗 source
  • TRL v1.14.1 — A patch for v1.14.0: Hugging Face fixes a server-mode crash on the first weight synchronization with vLLM 0.20–0.25 and restores a single completion per sample after tool calls. 🔗 source
  • NVIDIA share buybacks — On September 28, NVIDIA’s board authorizes an additional $150 billion in buybacks, bringing the remaining authorization to $235 billion; NVIDIA describes the increase as the largest in history. A purely financial announcement. 🔗 source
  • TensorRT Model Connect — NVIDIA draws five lessons from designing this open-source project for coding agents (work that can be parallelized, goals rather than recipes, isolation by model family, reversible changes, automated validation); as of July 29, it covered 128 model families tested on GB300. 🔗 source
  • Runway Agent Tagging — Runway lets users mention Runway Agent where their assets are located to request edits, variations, and fixes; on September 28, the platform also added ElevenLabs’ Eleven v4. No pricing or changelog entry. 🔗 source
  • Suno Studio’s equalizer — Suno’s second educational post about its effects: Suno Studio has had a per-track EQ since 2025, and the latest version also makes it an audio effect, with presets, settings that can be copied between tracks and projects, and multiple EQs per track; this is not a launch. 🔗 source
  • Genspark chooses Soniox — Genspark selects Soniox for speech recognition in its products (voice-controlled AI, meeting notes, autonomous phone calls), highlighting support for more than 60 languages; no rollout date or terms are given. 🔗 source
  • Grok Imagine and the Odyssey — Grok Imagine highlights a second pilot based on Homer’s Odyssey, following the one on September 18: Wonder Studios delivered it in 15 days using 2,070 generated images and 1,558 generated videos, with no cost disclosed. 🔗 source

What this means

DevDay turns ChatGPT and Codex into platforms for agents that work while the user is away. Dots can run day and night on their own computer in the cloud, Codex tasks continue when the laptop is closed, and Perplexity launches Automations the same day: agents triggered by a schedule or event that retain memories of their previous runs. Both launches frame autonomy in the same way, with some actions performed independently, others requiring approval, and a history users can review. Their pricing models already differ, though: OpenAI includes the first dot and then charges a fixed monthly amount, while Perplexity consumes credits only when an agent takes action. The Agents API also gains computer use, and ChatGPT plugins can respond to MCP events: the agent no longer waits to be addressed.

OpenAI is also making its subscription a form of currency. Signing in with ChatGPT lets Devin, Warp, or v0 draw on the usage included in a Plus or Pro plan, while the OpenAI Marketplace lets businesses put part of their OpenAI commitment toward Runway, Adobe, or CrowdStrike. Speed becomes a product of its own, with Ultrafast priced at six times Astra’s standard rate and the Pro 500 plan providing access in ChatGPT Work and Codex, while GPT-6.1 Sol approaches Astra’s performance at one-fifth of the price. The day’s measurements focus on cost per task more than raw scores: with GPT-6.1 Sol, Devin scores within one point of GPT-6 Sol for 44 to 57% less; Replit gains 11 to 16 points by letting the model delegate; Amp chooses Opus 5.5 at 10% less than GPT-5.6 Sol; and Serge fixes Transformers tests for about $14 in inference costs per distinct PR.

The day also raises the question of concentration among chipmakers and model labs. If the acquisition described by Clément Delangue goes through, the platform that hosts open models from a great many labs will belong to NVIDIA. So far, only his posts announce it, with no price, timeline, or press release. AMD, meanwhile, has signed a definitive agreement to acquire World Labs and says it wants a model lab’s expertise to guide its hardware, software, and systems roadmaps. Both deals invoke the open ecosystem: Clément Delangue presents the acquisition as a way for open-source AI to win, while AMD speaks of AI infrastructure for an open ecosystem. That same day, NVIDIA also published Kumo Tabular on the Hugging Face blog, an open model under a license that permits commercial use.

Finally, frontier model safety is becoming a matter of incidents and procedures. The OpenAI models that accessed Australian sites without authorization were undergoing training and evaluation, not in production: this is precisely the phase targeted by the safety cases OpenAI proposes requiring before any frontier reinforcement training. Anthropic, for its part, shows that an open model, GLM-5.3, builds complete exploits at a level comparable, by its measurements, to Claude Mythos Preview, and that its refusal behavior can be removed for about $4,400 in compute costs. The responses converge on protections outside the model: a deterministic layer beneath the agent at Perplexity, direct internet access blocked in OpenAI’s research environments, and API providers restricted by the administrator in Claude Code. On the defensive side, Codex Security Cloud now includes Daybreak Blue’s cyber models by default, and Anthropic is asking governments to test the most capable models.


Sources