Search

GPT-6 for everyone in ChatGPT, Claude Haiku 5.5 and RTX Spark PCs available for preorder

ai-powered-markdown-translator

Article translated from French to English with gpt-6.1-sol.

View project on GitHub ↗

GPT-6 and Intelligent UI are coming to all ChatGPT users: GPT-6 Sol for paid plans starting October 7, GPT-6 Luna for Free and Go starting October 8. Anthropic rounds out its 5.5 family with Claude Haiku 5.5, starting at 0.10 dollar per million input tokens, while Microsoft’s Windows event opens preorders for RTX Spark PCs and announces that GitHub Copilot will soon be able to delegate its tasks to a local model. OpenAI is also publishing 722 manuscripts of mathematical results produced by an internal model.


ChatGPT moves to GPT-6 and Intelligent UI for everyone, reads audio files and prepares College Planner

October 7 — OpenAI is extending GPT-6 to all ChatGPT users. This is not a new base model: GPT-6 Sol and GPT-6 Luna have already been available to paying customers since September (covered here on September 22). This time, October versions of these two models, tuned for everyday conversation, are coming to the Chat tab: GPT-6 Sol for Plus, Pro, Business and Enterprise starting October 7, GPT-6 Luna for Free and Go starting October 8. They replace GPT-5.6 Sol and GPT-5.6 Luna in ChatGPT, while Codex and ChatGPT Work remain on the September versions. OpenAI says this rollout serves the more than 1.2 billion people who use ChatGPT each week.

The most visible addition is called Intelligent UI. GPT-6 composes its responses with text, visuals and interactive elements (charts, buttons, forms, interactive diagrams, maps) and can create a small tool on request, such as a savings calculator, a bill splitter or a game; when plain text is enough, ChatGPT sticks to it. To do this, OpenAI uses a library of native components and a compiler that displays the interface as it is generated. Intelligent UI works at effort levels from Instant to Extra High, with no separate quota, but neither with Pro effort, handled by GPT-6 Pro (powered by GPT-6 Astra), nor in the older desktop applications for macOS and Windows.

A second change: ChatGPT can start responding while it is thinking or using tools, then complete its response without a new prompt. According to OpenAI’s internal evaluations, GPT-6 Instant starts responding 44 % sooner on average than GPT-5.6 Instant on questions that require a web search, and GPT-6 Extra High starts within the same timeframe as GPT-5.6 Medium, with a better overall score than GPT-5.6 Extra High.

Plan or effort settingModel used in ChatGPTSchedule and details
Plus, Pro, Business and EnterpriseGPT-6 Sol (Instant, Medium, High, Extra High)Worldwide rollout starting October 7
Free and GoGPT-6 LunaStarting October 8
Pro effort (Pro 100 and 200 dollars, Business, Enterprise)GPT-6 Pro, powered by GPT-6 AstraWithout Intelligent UI
ChatGPT Work and CodexSeptember versions of GPT-6 Sol and GPT-6 LunaNo change

On safety, the system card published the same day classifies these versions as having “High” capabilities in cybersecurity as well as biology and chemistry under the Preparedness Framework, without reaching that threshold in self-improvement; they retain GPT-5.6’s safeguards, with better resistance to circumvention attempts (jailbreaks), including across multiple turns. In the API, the chat-latest snapshot now points to this new ChatGPT model, with OpenAI recommending the GPT-6 family for production.

GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone. Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot. — @OpenAI on X

🔗 GPT-6 and Intelligent UI for everyone · System card, October update · GPT-6 and other models in ChatGPT

Audio files in ChatGPT

October 6 — ChatGPT now accepts audio files as attachments. You can attach a recording to a conversation to get a transcript, a summary or answers about its contents, or to turn a meeting, an interview or a lecture into structured notes, or even a draft follow-up email. The feature is limited to paid subscriptions and workspaces, including Enterprise: the Free plan does not have access “for now.” Supported formats are WAV, MP3, OGG, FLAC, AAC, M4A and PCM, as well as WebM and MP4 if they contain only audio, up to 512 MB per file. OpenAI warns that transcripts may contain errors, speaker identification remains unreliable and performance varies by language.

🔗 ChatGPT release notes · Uploading files and audio to ChatGPT

ChatGPT for Teens: first usage figures and upcoming College Planner

October 7 — OpenAI provides its first progress update on ChatGPT for Teens, the experience applied by default to accounts identified as belonging to users under 18. According to the company, nearly 1.2 million teens used Learning Visualizations in one week and more than 180 000 used Study Mode; they spend an average of less than 15 minutes a day on it, and fewer than 2 % exceed three consecutive hours. The newly announced College Planner will arrive “soon,” initially in the United States for students in grades 10 through 12 aiming for a four-year college: a single plan will bring together application requirements, deadlines, tasks and financial aid steps, including scholarships. OpenAI is also supporting College Advising Corps and, for three years, the Digital Wellness Lab at Boston Children’s Hospital, without disclosing an amount.

🔗 Helping teens learn, plan, and shape the future of AI


Claude Haiku 5.5, about 75 % cheaper than Haiku 4.5 according to Anthropic

October 7 — Fifteen days after Opus 5.5 and nine days after Sonnet 5.5, Anthropic completes the Claude 5.5 family with Claude Haiku 5.5 (claude-haiku-5-5). The model targets high-volume, cost-sensitive tasks (summaries, context compaction, database queries, classification), the role of coding sub-agent for Opus 5.5 and Sonnet 5.5, and uses where speed matters, such as live customer support or browser control: it is Anthropic’s fastest model at standard speed, with Opus models in Fast Mode remaining faster. It is also the first Haiku with an effort setting (medium by default), with a context of 1 million tokens and up to 128 000 output tokens.

Pricing has two tiers. Up to 100 000 prompt tokens, Haiku 5.5 costs 0.10 dollar per million input tokens and 0.50 dollar for output, 90 % less than Haiku 4.5; above that, it rises to 0.50 and 2.50 dollars, 50 % less. The advertised savings of “about 75 %” are an average calculated by Anthropic: 90 % of requests sent to Haiku 4.5 stayed below the threshold, and the new tokenizer counts about 30 % more tokens for the same text.

Pricing type (per million tokens)Haiku 5.5 up to 100 000 prompt tokensHaiku 5.5 above 100 000 tokensHaiku 4.5
Input tokens0,10 dollar0,50 dollar1 dollar
Output tokens0,50 dollar2,50 dollars5 dollars
Cache reads0,01 dollar0,05 dollar0,10 dollar
Cache writes0,125 dollar0,625 dollar1,25 dollar

On the benchmarks published by Anthropic, the gap with Haiku 4.5 is wide, and Haiku 5.5 beats OpenAI’s GPT-6 Luna everywhere both appear. It remains behind Sonnet 5.5, however, which Anthropic still recommends, along with Opus 5.5, for complex agentic coding.

Benchmark evaluated (Anthropic figures)Haiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
Terminal-Bench 4.0 (agentic coding)39,2 %0,0 %16,4 %70,6 %
OSWorld 2.1, offline subset72,4 %15,7 %48,9 %83,9 %
Humanity’s Last Exam, without tools45,9 %10,2 %—56,9 %
FrontierCode 1.1, Main46,4 %—42,4 %52,1 % (Xhigh effort)
Chartography, without tools46,4 %6,4 %29,1 %61,6 %

Six customers shared accounts of their early trials: HubSpot reports 92.8 % on its CRM suite, its best score to date, AlphaSense 0.84 versus 0.76 for Haiku 4.5 across 400 queries, and Box an 11-point improvement with roughly half the latency. Migrating from Haiku 4.5 takes work: the guide lists eleven changes, including the removal of budget_tokens, the temperature, top_p and top_k parameters, and prefilling, and Claude Code’s /claude-api migrate command automates part of it. On safety, Anthropic describes cyber safeguards that are more restrictive than Haiku 4.5’s but slightly less restrictive than those of its other recent models: more defensive tasks are allowed than with Sonnet 5.5, while penetration testing remains blocked.

The launch comes with a price cut: Sonnet 5.5 cache reads drop from 0.20 to 0.10 dollar per million tokens starting October 7, making the model about 20 % cheaper for most agentic tasks, according to Anthropic. Haiku 5.5 is available now through the Claude API, Amazon Web Services, Google Cloud and Microsoft Azure, as well as in Claude Code.

Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released. On average, it costs around 75% less to run than Claude Haiku 4.5. — @claudeai on X

🔗 Claude Haiku 5.5 announcement · Haiku 5.5 migration guide

Haiku 5.5 in Cursor: 48.4 % on CursorBench

October 7 — Cursor makes Claude Haiku 5.5 available on launch day, ready to enable in its model settings. According to Cursor, it costs 10 times less than Claude Haiku 4.5 for short requests. On CursorBench, the ranking Cursor publishes on its website, Haiku 5.5 at Max effort places 10th with a 48.4 % success rate at 1.12 dollar per task, just ahead of Sonnet 5.5 at High effort, whose costs Cursor recalculated on October 7 to account for its new pricing, and ahead of Opus 5 at Max effort. It does more work, however: an average of 325 934 tokens and 163 steps per task, compared with 37 391 tokens and 41 steps for Sonnet 5.5 at High effort.

Configuration tested by CursorRankingCursorBench scoreCost per task
Opus 5.5, Max effort157,8 %13,43 dollars
Haiku 5.5, Max effort1048,4 %1,12 dollar
Sonnet 5.5, High effort1147,8 %1,20 dollar
Opus 5, Max effort1346,6 %11,95 dollars
Haiku 5.5, Low effort5430,9 %0,08 dollar

🔗 Cursor announcement on X · CursorBench ranking


A monthly API credit of 100 to 500 dollars for Max and Team plans

October 7 — Announced in the same post as Haiku 5.5, a monthly credit usable on the Claude Platform is coming to Max and Team subscribers, with a rollout spread over a few days. It can be used to build your own tools, applications and agents with the API; according to @ClaudeDevs, it works with all models, including Haiku 5.5, both in your own code and in third-party tools.

Eligible planMonthly API credit
Max 5x100 dollars
Max 20x200 dollars
Team, Standard seat20 dollars per seat
Team, Premium seat100 dollars per seat
Cap for a Team workspace500 dollars, pooled
Free, Pro and EnterpriseNot eligible

On Team, seat credits are added to a shared pool: three Standard seats and two Premium seats provide, for example, 260 dollars per month. The credit covers all Claude API models, including the Message Batches API, as well as the Console Playground, Claude Managed Agents and the Claude Agent SDK. It does not cover interactive Claude Code usage (terminal, IDE, desktop application or web), extra usage beyond plan limits, or Claude models served by Amazon Bedrock, Google Cloud Vertex AI or Microsoft Foundry, and it does not change the usage limits for Claude, Claude Code or Cowork.

To claim it, you need a subscription that has been active for at least seven days on an eligible plan, then link a Claude Console organization from the billing settings on claude.ai, without adding a payment method. Only one organization can be linked, and only support can change it. The credit renews each billing cycle, does not roll over and is used before purchased credits; once it runs out, calls stop until the next credit arrives, unless the organization has purchased credits or enabled automatic top-ups, and nothing is charged to the Claude subscription. In May, Anthropic announced a monthly credit for programmatic usage starting June 15; the help page clarifies that the new credit is separate from those “Agent SDK credits,” which it says are unavailable.

🔗 Help page: monthly API credits for Max and Team plans · @ClaudeDevs announcement


Local AI on Windows: NVIDIA opens preorders for RTX Spark PCs and previews DGX Station for Windows

October 7 — NVIDIA and Microsoft are using the Windows AI and Surface event, held in San Francisco with a conversation between Jensen Huang and Satya Nadella, to commercially launch RTX Spark, the Windows PC chip announced at Build on June 2. Laptop preorders have been open since October 7, with availability on October 16; mini-PCs will follow in November. Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte are preparing machines, and Microsoft is introducing a Surface Laptop Ultra built around the chip. NVIDIA’s post does not list any prices.

RTX Spark combines a Blackwell RTX GPU and a Grace CPU connected at 600 Go/s, delivering one petaflop of AI compute in FP4 and up to 128 Go of unified memory. NVIDIA highlights running large models locally, without sending data to the cloud or metering usage, with the same CUDA stack as on its servers. The chip also targets creators (NVFP4, AV1 encoding and hardware 4:2:2) and gamers, with games running at 1440p above 100 frames per second thanks to DLSS 5.

For businesses, NVIDIA is previewing DGX Station for Windows, also announced at Build in June, which it describes as the Windows ecosystem’s first desktop AI supercomputer: enough to run models with around a trillion parameters locally without leaving Windows, while Linux tools remain accessible through WSL. There is no date or price, only a signup to receive notifications. Meanwhile, Microsoft is making Microsoft Execution Containers (MXC), the infrastructure that runs agents in the background under Windows’ control, generally available.

Item comparedRTX SparkDGX Station for Windows
ChipBlackwell RTX GPU (up to 6 144 cores) and Grace CPU (up to 20 cores)GB300 Grace Blackwell Ultra Desktop superchip
MemoryUp to 128 Go, unified748 Go, coherent
AI compute in FP41 petaflopUp to 20 petaflops
Announced availabilityLaptops available for preorder, available October 16; mini-PCs in NovemberPreview, with no date or price

🔗 NVIDIA post on the Windows event


Developing locally on Windows: Copilot will route to MAI Code 1.1 Flash, its sandbox becomes generally available, Replit prepares a desktop app

October 7 — GitHub Copilot will soon be able to choose on its own between a model running on the developer’s machine and a model in the cloud. According to the Microsoft and GitHub post published on the Microsoft Command Line blog, this routing is coming “by the end of the month”: it is not available yet. GitHub presents it as the next step for Project HydraFusion, its orchestrator that already distributes tasks across several models, and highlights savings in AI credits without quantifying them.

Two ways to use it are planned in Copilot CLI, the GitHub Copilot app and VS Code. Auto mode will choose, turn by turn, between local and cloud inference based on the task context and cache state; developers will also be able to select a local model themselves, either MAI Code 1.1 Flash through the Windows ML provider or any model exposed through a local access point (endpoint) compatible with the OpenAI API. The post reiterates that local inference does not make the session offline.

The demonstration runs on a Surface Laptop Ultra equipped with RTX Spark. Microsoft AI is running a local version of MAI Code 1.1 Flash on it, a mixture of experts (mixture-of-experts) specialized in code, with 137 billion parameters, of which 6,8 billion are active. Quantized to around 3,3 bits per weight, it weighs in at 53 Go, 80 % less than the cloud variant, with peak memory usage of 75,5 Go at 256 000 tokens of context.

Benchmark evaluated (Microsoft tests on October 5)MAI Code 1.1 FlashQuantized version on the deviceGPT OSS 120B (Unsloth’s GGUF version)
SWE-Bench Verified (500 tasks)72,6 %70,80 %32,0 %
Terminal-Bench 2.1 (89 tasks)62,9 %66,29 %23,6 %

Microsoft specifies that these tests ran on a llama.cpp runtime for Windows ARM64 and that results vary by device and configuration. A first component is already available in the terminal: since Copilot CLI preview version 1.0.94-0, the /model command discovers models from a local Ollama instance. Nothing is added without confirmation, Ollama and the model must already be installed, and only models that support tool calling and continuous output (streaming) are offered.

🔗 Bringing local models and sandboxed tools to Windows and GitHub Copilot · Discover local models in GitHub Copilot CLI

Copilot’s local sandbox becomes generally available

October 7 — GitHub Copilot’s local sandbox is leaving the public preview opened on June 2: it is becoming generally available in Copilot CLI, the GitHub Copilot app and VS Code sessions that use Agent Host, at no additional cost. Tools and commands launched by Copilot then run with restricted access to files, the network, Git and GitHub CLI credentials, and other system capabilities, according to policies set by the developer or their organization; a company can enforce settings that developers cannot relax, regardless of the model used. The engine, Microsoft eXecution Container (MXC), is an open source library from the Windows team that relies on ProcessContainer on Windows, Seatbelt on macOS and bubblewrap on Linux, without a virtual machine. Its limitations are documented: built-in file tools are controlled by Copilot itself rather than by the system, and remote MCP servers remain outside the local sandbox.

🔗 Local sandboxing for GitHub Copilot now generally available

Replit builds applications locally on Windows

October 7 — Replit, together with Microsoft, announces the preview of a desktop application that builds and runs applications directly on the user’s computer, on Windows. Replit already offered a desktop application: the new feature is this local building process, where each build runs in its own sandbox, based on Microsoft Execution Containers and OpenShell, NVIDIA’s open source runtime. Early access is through a waiting list, with no date or pricing, and no macOS version is mentioned. Replit presented this preview onstage during the Windows event on October 7.

🔗 Replit’s announcement on X · Preview waiting list


OpenAI publishes 722 manuscripts of mathematical results produced by an internal model

October 6 — OpenAI is publishing a large collection of mathematical results all at once, produced by an internal frontier model that is not publicly available. Announced on the evening of October 6, the openai/math GitHub repository brings together 722 manuscripts grouped into 372 families organized by discipline: a main result may be accompanied by supporting arguments, consequences or alternative proofs.

The vast majority of the results come from the same procedure: around 4 000 problems posed to the model, with an average equivalent of three hours of ChatGPT Pro thinking per result. OpenAI explains that it expanded its evaluations to open research problems after saturating its existing mathematical evaluations. Two works are exceptions to this procedure: a zero-free region for the Riemann zeta function, whose wording was edited by a human for readability, and the proof of a special case of the Hodge conjecture, that of CM abelian varieties.

Metric published by OpenAIAnnounced value
Manuscripts published722
Families of results372
Problems posed to the modelAround 4 000
Average compute per resultAround three hours of ChatGPT Pro thinking
Reasoning summaries published10

Not all results are verified in the same way. Many proofs are formalized in Lean, a language that allows a proof to be checked by computer, but not all: OpenAI warns that “some non-formalized results may contain errors,” promises to correct them quickly and will add new formalizations as they become available. The ten summaries of the model’s reasoning cover topics such as the irrationality exponent of π, Mahler’s conjectures, Kaplansky’s direct finiteness conjecture in characteristic two, and the isomorphism of free group factors.

The publication itself was prepared with the Institute for Advanced Study’s independent advisory group on mathematics and AI: review and citation protocols, corrections published as new versions, and history preserved. OpenAI also announces funding for workshops, conferences and programs devoted to these results, and says it is working on “responsible” access for scientists to the model that produced them, without a timeline.

We’re releasing a broad range of new mathematical results produced by an internal frontier model. We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results. — @OpenAI on X

🔗 Sharing AI progress in mathematics · openai/math repository on GitHub


Perplexity releases pplx-embed-v2-late, multivector embeddings for text and images

October 7 — Perplexity Research releases pplx-embed-v2-late, two late-interaction (late interaction) embedding models with 0.6B and 9B parameters, available on Hugging Face under the MIT license. Where a dense embedding compresses each document into a single vector, these ColBERT-style models retain a 128-dimensional vector per token and score each query-document pair using MaxSim: each query token is matched to the closest token in the document. They search text, images and rendered pages (PDFs, slides, scans) without using OCR, preserving tables, figures and page layout.

Both sizes share the same embedding space because they are distilled token by token from an internal 18B teacher derived from a 27B base. A corpus indexed with the 9B can therefore be searched using queries encoded with the 0.6B: on image-based ViDoRe v3, this configuration reaches 63,5 %, compared with 62,3 % using the 0.6B on both sides, with no additional query cost. The 0.6B, pruned from Qwen3.5-0.8B, thus serves as a lightweight query encoder, even on the device.

Benchmark evaluated (Perplexity measurements)9B score0.6B scoreComparison point
72 specialized tasks (nDCG@10)81,3 %78,0 %The 9B leads the next model by 1,6 points
Q2D-Web, web search (Recall@1000)74,8 %73,6 %nemotron-embed-8b: 69,3 %
Image-based ViDoRe v3 (nDCG@10)65,2 %62,3 %Only EVIE outperforms the 9B
BrowseComp+ (GPT-OSS-120B agent)64,0 %—Next-best ColBERT: 4,9 points lower
MADQA (Gemini 3.5 Flash agent)92,4 %90,1 %Mixedbread Agentic Search: 93,4 %

The 9B does not win everywhere, and Perplexity says so: Tencent’s EVIE outperforms it on image-based ViDoRe v3, gemini-embedding-2 on MIRACL-Vision and the internal PPLX-Q2I benchmark, and Mixedbread Agentic Search achieves a higher score on MADQA, a gap that Perplexity considers to be within its confidence interval. The company also acknowledges that storage and scoring costs increase with document length. A technical report is announced for “later this year,” and Perplexity plans to gradually make its late-interaction, dense and contextual embeddings available on its API platform, without a date.

🔗 Perplexity Research post · pplx-embed-v2-late-9b on Hugging Face


Coding agents: Claude Code 2.1.293, Codex CLI 0.161.0, Cursor on iOS, Antigravity 2.21.0 and Warp Factories

Five coding agent tools are evolving: Claude Code adopts Haiku 5.5, Codex CLI connects to MCP servers from the terminal, Cursor can be controlled from an iPhone, Antigravity groups its agent’s actions, and Warp adds support for Microsoft’s tools.

Claude Code 2.1.293 switches to Haiku 5.5 by default

October 7 — Released a few minutes after the Haiku 5.5 announcement, Claude Code version 2.1.293 makes it the default Haiku model on the Anthropic API. It adds an agentType field to the subagentStatusLine payload to distinguish custom subagents, and an isDeferred option that exposes the schema of tools declared by mods from the outset. The bulk of the release consists of 38 fixes out of 56 entries: after compaction, Claude could place its last pre-compaction actions after the compaction, then remove or redo completed work; path-scoped rules and nested CLAUDE.md files also load when Claude reads a file with cat or grep in Bash; a memory leak in HTTP MCP connections is eliminated. Two recent changes are reverted (the auto-mode refusal message from 2.1.281 and a cloud-session fix from 2.1.290), and Claude Tag’s channel rule limit increases from 20 to 50.

🔗 Claude Code 2.1.293 on GitHub

Codex CLI 0.161.0 gates Daybreak behind an option

October 7 — Codex CLI 0.161.0 is the first stable release since 0.160.1 on October 5. The /mcp login <name> command allows users to connect to an MCP server without leaving the current terminal session, and voice conversations gain a choice of microphone, speaker and input channels. Daybreak, OpenAI’s cyber lineup, moves behind an option: the /daybreak toggle and Daybreak state in the status line remain hidden until the user enables it with --enable cli_daybreak (or features.cli_daybreak=true), and without it, automatic Cyber routing is omitted. For a single turn, codex exec --cyber-access-program selects a Cyber access program. On Amazon Bedrock, multi-agent V2 and Ultra reasoning effort arrive on compatible models, and Bedrock Mantle accepts AWS GovCloud regions. Fixes address permissions, thread resumption and detection of a corrupted SQLite database.

🔗 Codex CLI 0.161.0 on GitHub

Cursor controls agents on your computer from iOS

October 6 — Cursor’s iOS app now displays agents running locally on your computer: you can see what each one is doing and reply to it, and the announcement tweet also mentions launching new tasks. The agents stay on the machine, and the app connects to it: the computer must therefore remain powered on and online, and a Keep this computer awake setting prevents it from sleeping while plugged in with the lid open. The feature is enabled by default, except in Enterprise organizations, where an administrator must enable it; pairing is confirmed in the desktop app, and cloud agents are not required. The iOS app released in late June focused on cloud agents: local agents are now becoming accessible from your phone.

🔗 Cursor changelog

Antigravity 2.21.0 groups its agent’s actions

October 6 — Google’s Antigravity app moves to version 2.21.0, with 9 improvements and 14 fixes. Advanced search finds conversations by their content, using ⌘+K (Ctrl+K on Windows). The scheduled tasks tab (Scheduled Tasks) becomes Automations: you can run an automation immediately with Run Now, or ask the agent to guide you through creating a new one with Create with Prompt. File edits, terminal commands and reads performed by the agent between responses are now grouped into a single collapsible line, accompanied by a short summary. Among the fixes, a truncated MP4 or MOV video read by the agent could cause the rest of the conversation to fail, and a settings file saved with Windows Notepad or PowerShell would stop working. The changelog warns that a release may take a few days to reach all users.

🔗 Antigravity changelog

Warp Factories adds support for Microsoft Teams and Azure DevOps

October 7 — Warp is extending Warp Factories, its software factory platform (software factories), to Microsoft Teams and Azure DevOps Services. In Teams, mentioning the Warp app in a configured channel or thread assigns the task to the factory with the conversation’s context; it reports its progress in the same thread and posts its pull requests there for review. In Azure DevOps, you can mention the factory from a work item (work item) or a pull request, or trigger runs on their events; each factory receives its own identity for Git operations, with access limited to selected repositories. Both integrations are available to all Warp Factories customers. Warp presents this dual native support as an industry first; Devin already integrated with Teams and Azure DevOps Server.

🔗 Warp blog post


APIs and platforms: computer use and browser use toolsets in the Claude SDKs, Grok 4.7 on Microsoft Foundry

Two announcements for developers building on hosted models: Anthropic simplifies computer and browser control, and Grok 4.7 reaches general availability at Microsoft.

Computer use and browser use toolsets in the Claude SDKs

October 7 — Claude’s Python and TypeScript SDKs now include computer use and browser use toolsets (toolsets) in beta. Until now, the API indicated what Claude wanted to click or type, and developers had to write their own loop to translate each action into a command. The SDK now handles this loop: the developer subclasses BetaAbstractBrowserToolset20260801 or BetaAbstractComputerToolset20260801, writes one method per action, and the SDK routes calls, applies the supplied URL and file policies, requests approvals and builds the results returned to the model. Anthropic provides neither a browser nor a ready-to-use driver: you connect your own automation or a driver from Browser Use, Browserbase, E2B or Daytona, and a minimal Chrome DevTools Protocol example is published in the claude-quickstarts repository. JavaScript execution and file uploads are disabled by default, and without a URL policy, no addresses are checked.

🔗 @ClaudeDevs thread on X · SDK toolsets documentation

Grok 4.7 reaches general availability on Microsoft Foundry

October 7 — Grok 4.7, the SpaceXAI model released on September 21, is available on Microsoft Foundry, @SpaceXAI announced overnight. Microsoft’s documentation lists it as generally available, among the models sold directly by Azure: text and image input, a context of 500,000 tokens (input and output combined), tool calling, access through Chat Completions or the Responses API, and reasoning effort from low to xhigh, with high as the default. Microsoft commits to offering Grok models in general availability, starting with Grok 4.7, for at least six months; Grok 4.6 is still listed in preview. The Azure pricing page did not yet list Grok 4.7 on the evening of October 7. Following Amazon Bedrock (September 28) and Google Cloud in preview (October 1), Grok 4.7 is now offered by all three major cloud providers.

🔗 Deploy and use Grok models in Microsoft Foundry


Open models: Liquid AI’s Open d1, Ai2’s Bolmo in Nature with Bwen 8B and Blama 8B

Two releases with open weights, one for making fast decisions at the network edge, the other for reading text byte by byte.

Open d1, Liquid AI’s open decision models

October 7 — Liquid AI opens up its family of decision models with Open d1: two models with open weights, d1-3B (text and image) and d1-omni-600M (text with image or audio), the latter as an early research release. They do not generate text: they assign a probability to each permitted answer in a single pass. According to Liquid AI, d1-3B scores 48.57 on Decision Index 0.2.1, the best score below 10 billion parameters, ahead of Decider 35B-A3B (47.11), and averages 82.9 across seven public datasets. Measured with NVIDIA, its latency per question ranges from 8 ms on an RTX 4090 to 50 ms on a Jetson Orin Nano, making it suitable for the network edge (edge). No vision or audio benchmarks are published. The weights are on Hugging Face under Liquid AI’s own license (lfm1.0), with GGUF versions; the d1 released on September 29 was accessible only through an API.

🔗 Liquid AI blog post on Hugging Face

Ai2’s Bolmo in Nature, with Bwen 8B and Blama 8B

October 7 — Ai2 publishes in Nature, online on October 7, the method behind Bolmo, its family of fully open models that read bytes directly rather than subwords. Reading bytes avoids the blind spots of a fixed vocabulary (spelling, unusual strings, certain writing systems), but until now required full training from scratch. Byte conversion (byteifying) instead starts with an already capable subword model and transforms it through a short period of additional training. Bolmo 1B and 7B, derived from Olmo, date back to December; Ai2 also presents Bwen 8B (derived from Qwen 3 8B, Apache-2.0 license) and Blama 8B (derived from Llama 3 8B, llama3 license), whose repositories have existed since August 26, along with “Stage 1” checkpoints where only the new byte layer is trained. According to Ai2, the two models approach their original models’ performance and Bwen 8B outperforms Bolmo 7B, with no figures published.

🔗 Ai2 blog post


SynthID Detector opens to everyone for checking images, videos and audio

October 7 — Google opens SynthID Detector to everyone, the tool that detects invisible SynthID watermarks in AI-generated content. Introduced last year in a preliminary version for media professionals, it is available worldwide and in English on synthid.com. Users upload an image, a video or an audio file. Verification covers content from Google and its partners OpenAI, NVIDIA and Kakao, with Apple coming soon. The portal warns that it is not a general-purpose AI detector: content from a model without SynthID, or content that has been heavily modified, may return “not detected.” Google reports more than 180 billion images and videos and 240,000 years of audio watermarked since 2023, compared with the 100 billion and 60,000 years announced in July, and more than a million checks per day in Search, the Gemini app and Chrome.

🔗 Google expands SynthID Detector for AI content


Security: GitHub’s classifier for leaked secrets

October 7 — GitHub deploys a fine-tuned ModernBERT classifier dedicated to leaked secrets and built with Microsoft Applied Sciences: it reads the code surrounding a value to identify likely credentials, including passwords without a recognizable format. It evaluates a batch of candidates in under 2 ms and could, according to GitHub, “more than double” the secrets blocked at push time. Starting now, alerts for AI-detected passwords switch to this model at no additional cost. AI push protection, in private preview, is due to arrive “later this month” for eligible organizations, and checks through Copilot’s /security-review command will enter private preview “soon,” both using AI credits. The essay by Erin Havens, GitHub’s security product lead, notes that one in three pull requests involves an AI agent and that push protection stops only about 30% of newly detected secrets.

🔗 Secret protection must scale with software


Infrastructure: GitHub rebuilds Git for agent activity

October 6 — GitHub is rebuilding its Git infrastructure to keep pace with agentic development. According to Brian Celenza’s engineering blog post, total Git activity grew from 218.2 to 473.3 billion events per month between September 2025 and August 2026; developers and agents produced 7.38 billion commits in September 2026, more than five times as many as a year earlier, and pushes increased by a factor of 4.9. The current architecture, Spokes, replicates each repository across several servers and validates each update by majority vote: adding copies for reads therefore slows down writes. The new architecture separates storage and compute, with authoritative data in Azure Blob Storage and lightweight processes (workers) serving reads from a cache. In its internal benchmarks, GitHub measures up to 35 times the write throughput, with no switchover date.

🔗 Building Git infrastructure for agent-scale development


Voice AI: ElevenLabs signs a memorandum of understanding with an Indian state

October 7 — At its Bengaluru Summit, ElevenLabs signed its first memorandum of understanding (MOU) with an Indian state government for a voice AI pilot. The company does not name the state or provide a timeline, scope or amount: the announcement consists of a single tweet, with no blog post. The accompanying figure sheds light on the stakes: over the past year, ElevenAgents has conducted more than 100 million conversations in India, more than 70% of them in Hindi, Kannada, Tamil and Telugu. The Summit brought together more than 500 business leaders, developers and partners.

🔗 @ElevenLabs announcement on X


Research: PivotOPD, NVIDIA’s method for recovering from an agent’s decisive error

October 7 — NVIDIA researchers present PivotOPD, a training method for agents that perform multiple rounds of actions. Their finding: on ALFWorld, 59% of failures across three Qwen3 models contain a “pivot error,” made early, after which the agent continues in the wrong direction. PivotOPD extends on-policy distillation (on-policy distillation): at each pivot error, a teacher model indicates the correct action, then a recovery action for subsequent turns, so the student learns to avoid the error and recover from it. Against 13 baseline methods, it achieves the best average on ALFWorld, WebShop and Search-based QA with Qwen3 students of 1.7B and 8B, and recovers from 72.7% of the 72 replayed pivot errors, compared with 20.3% for standard distillation. The improvement carries over to coding: a Nemotron-3.5 student rises from 62.8% to 66.0% on SWE-Bench Verified. The paper is on arXiv; the code is announced as coming soon.

🔗 PivotOPD project page


Optimization: mPDLP distributes giant linear programs across multiple GPUs in cuOpt

October 7 — NVIDIA adds mPDLP to cuOpt, its open source optimization library: a linear programming solver distributed across multiple GPUs connected by NVLink, for large planning problems (supply chains, power grids). By grouping interdependent variables and constraints on the same GPU, it limits communication and reduces peak memory per GPU by up to a factor of 6. On a DGX B200 node, speedup reaches 11.4 times for PDLP computation and 4.2 times end to end; compared with D-PDLP, mPDLP is 1.2 to 2.5 times faster on most large problems, but slower on the three largest instances and on small problems. Kinaxis solves a supply chain with more than 135 million variables 3.3 times faster on eight H100 GPUs, and PSR solves an energy model with 185 million variables more than 5 times faster on eight B200 GPUs.

🔗 NVIDIA blog post on mPDLP


Robotics: robots assemble GB300 test trays

October 7 — NVIDIA’s Seattle Robotics Lab explains how it is teaching robots to assemble GB300 test trays, which check modules before shipping. Two operations were selected with Foxconn, which manufactures these systems: placing and fastening a busbar with screws at 16 points, and plugging in four connectors mounted on flexible cables. The requirement set with Foxconn is a 99.5% success rate, taking at most twice as long as a skilled operator, with no collisions. The robots come close without reaching it: more than 95% success in 160 seconds for the busbar, against a target of 124, and 90 to 95% for the connectors, at 40 seconds per cable. The team combines a conventional perception and control pipeline, DOPER pose estimation and reinforcement learning in Isaac Lab, refined on the real robot. NVIDIA promises to release DOPER and its TALOS orchestrator, with no date.

🔗 NVIDIA technical blog post on GB300 trays


News in brief

  • ChatGPT iOS app 1.2026.272 — It displays page previews directly in responses, opens Codex task links in the app and, on the Codex Mobile side, redesigns the approval selector and speeds up streaming in long threads. 🔗 source
  • Radisson Hotel Group — According to an OpenAI case study, its ChatGPT plugin, built in six weeks with Accenture Song, converted about 1.5 times better than its organic search traffic in July–August 2026, and 54% of the payment and booking events from its ChatGPT advertising campaign are attributed to ad views alone. 🔗 source
  • Jump Trading — The quantitative trading firm assigns GPT-6 Astra agents analyses that can run for several days, drawing on numerous data sources, with human review at the end; OpenAI’s case study publishes no figures on gains. 🔗 source
  • ChatGPT plugin extensions — In the thread accompanying its video tutorial, OpenAI Developers names Adobe, Canva, Figma, Shopify, Instacart and Atlassian among the companies using them, alongside tldraw and MagicPath; introduced on September 29, they launch a plugin from the sidebar, handle @ mentions and add file viewers. 🔗 source
  • Every and Claude Managed Agents — Every’s team built a company agent on Claude Managed Agents that everyone uses in Slack to share their skills with each new model, then made it available to subscribers; Anthropic shares a video interview, with no figures. 🔗 source
  • Zed 1.23 and preview 1.24.1 — Version 1.23 becomes stable with its JSONL and NDJSON file previews, while preview 1.24.1 lets users drag a terminal tab into the Agent Panel, supports custom Amazon Bedrock models with tools, images and thinking, creates Git tags and reduces the Linux binary’s size by about 23%. 🔗 source
  • Puck, Amp’s meta-agent — It now supports multiple separate conversations: the + button opens one, clicking its name lets users switch between them, and conversations inactive for three days are filed separately. 🔗 source
  • Copilot CLI 1.0.93 — Now stable, it applies MCP server configuration changes between turns without restarting the session, and makes command sandboxing available to everyone through /sandbox and —sandbox; Copilot SDK 1.0.17 also becomes stable. 🔗 source
  • Copilot usage metrics — They have been undercounting agent activity since several IDEs began routing their sessions through the Copilot SDK; the fix requires an update (VS Code 1.139.0 now, Visual Studio 18.12 in October, JetBrains in late October, Eclipse and Xcode by November), and the lost data will not be recovered. 🔗 source
  • Gemini CLI v0.65.0-nightly.20261007 — This nightly release starts the 0.65.0 series with five fixes, including two addressing data loss: project settings become read-only in an untrusted folder, where gemini mcp add could empty .gemini/settings.json, and immediately exiting a resumed session no longer deletes its history. 🔗 source
  • Transformers.js 4.3.1 — The day after EmbeddingGemma 2 launched, the library computes its multimodal embeddings (740 million parameters, 768 dimensions) directly in the browser, using WebGPU or WASM, or in Node.js. 🔗 source
  • RL Environment Explorer — Adithya S K of Hugging Face presents this FineEnvs Space for browsing, visualizing and running the Hub’s reinforcement learning environments (OpenEnv, Harbor, MiMo, Verifiers, NeMo Gym): more than 12 million task entries, mostly from a few large OpenEnv environments. 🔗 source
  • Seb-9B — Stas Veretennikov releases this local decision model under the Apache-2.0 license (text, JSON or image, up to 20 options), which averages 0.792 across 17 public suites, compared with 0.786 for Jev 1.13; the author conducted all measurements and states that its training data included the training split of one benchmark, without its test cases. 🔗 source
  • Nemotron at the 2026 Olympiads — NVIDIA details the shared recipe (supervised fine-tuning, reinforcement learning, a loop that generates, verifies and corrects) behind a score of 535.4 out of 600 at IOI 2026, as an unofficial participant, and 30 out of 42 at IMO 2026, where the gold threshold was 29, and releases the 200-problem Nemotron-IMO-Bench benchmark. 🔗 source
  • Instadocs: AI Gone Wild — Netflix announces an October 12 release for this documentary reconstructing the cyberattack Hugging Face suffered in July, carried out, according to Netflix, by autonomous agents created by OpenAI researchers. 🔗 source
  • Runway in ChatGPT’s app directory — Once the plugin is installed and the Runway account connected, an @Runway mention in a conversation assigns it an image or video generation task; Runway specifies neither pricing nor credit costs. 🔗 source
  • Luma’s Face Swap — The feature replaces a character’s face in an existing shot to allow iteration without starting over; the announcement consists of two tweets, with no product page, pricing or model details. 🔗 source
  • DSX Air — NVIDIA shows how to test changes to an AI factory using this digital twin: agents apply the change, check configuration, security and isolation, then deliver a report, with deployment to production still subject to human approval; a methodology post, with no figures. 🔗 source
  • cuPhoton — An NVIDIA tutorial applies this open source toolkit to astronomical image analysis on GPUs, from reading FITS files to reviewing candidates; the reported speedups, up to 14,900 times, apply to certain operations rather than end-to-end processing. 🔗 source
  • Cosmos and SynthID — Content generated by NVIDIA Cosmos models on build.nvidia.com carries the SynthID watermark and can now be checked by anyone using the detector Google has made publicly available. 🔗 source
  • SpaceXAI TypeScript SDK 0.2.3 — It adds a fast service tier, interchangeable with priority, and the identifier grok-imagine-video-1.5-lite, a lightweight video model absent from the release notes: according to the data on the models page, it does not accept audio input and costs four times less per second than grok-imagine-video-1.5 at 480p. 🔗 source
  • RAG vs. LLM guide — Perplexity publishes an educational guide presenting RAG and LLMs as complementary, distinguishing three types of RAG (naive, modular, advanced) and explaining when an LLM alone is sufficient; no new features or figures. 🔗 source

What this means

Small models are becoming the mass-market product. Starting October 8, GPT-6 Luna answers free ChatGPT users, and Anthropic prices Haiku 5.5 at $0.10 per million input tokens for prompts under 100,000 tokens, while halving the price of Sonnet 5.5 cache reads. These models handle the bulk of the volume: everyday conversations, sub-agents, summaries, compactions. Cursor’s ranking is a reminder, however, that the price per token does not tell the whole story: at Max effort, Haiku 5.5 consumes nearly nine times as many tokens per task as Sonnet 5.5 at High effort, for a barely lower cost per task ($1.12 versus $1.20). Meanwhile, the monthly API credit included in Max and Team plans encourages subscribers to try these models in their own code.

AI is also moving onto the desktop. RTX Spark laptops arrive on October 16, DGX Station for Windows promises up to 20 petaflops on a desk, and Copilot is expected to assign some tasks itself to MAI Code 1.1 Flash locally by the end of the month. Replit now builds applications on the user’s machine. Agents executing code on a personal computer raise a security question, and the same building block keeps appearing: Microsoft Execution Containers (MXC), generally available on Windows, isolates both Copilot commands and Replit builds. GitHub, meanwhile, is deploying a classifier dedicated to leaked secrets and rebuilding its Git infrastructure, because one in three pull requests already involves an agent and commits have increased more than fivefold in a year.

The response is also becoming an interface. With Intelligent UI, ChatGPT responds with charts, forms or a small tool built on demand, and starts answering before it has finished thinking. For developers, Claude’s SDKs support the computer use and browser use loops, and Cursor lets users monitor agents working on their computer from an iPhone and reply to them. In all three cases, text is now just one part of the exchange: AI displays, clicks and reports back, while the user supervises.

Finally, many of today’s figures are measured by those publishing them: Haiku 5.5 benchmarks by Anthropic, MAI Code 1.1 Flash benchmarks by Microsoft, Q2D-Web by Perplexity, which designed it, Open d1 scores by Liquid AI, and the 35-fold increase in write throughput from GitHub’s internal tests. OpenAI itself warns that results in its 722 manuscripts that have not been formalized may contain errors, and its GPT-6 speed gains come from internal evaluations. These measurements remain useful for positioning products relative to one another; external evaluations, such as the CursorBench ranking, where a tool developer measures another company’s model, will help cross-check them.


Sources