Search

Anthropic opens Claude Code to mods, GitHub Copilot controls desktop apps, Black Forest Labs launches FLUX 3 Image

ai-powered-markdown-translator

Article translated from fr to en with gpt-6.1-sol.

View project on GitHub ↗

On Thursday, October 1, Anthropic launches mods, small TypeScript functions that hook into Claude Code events to rewrite its behavior and interface, shipping with version 2.1.287. On the same day, GitHub opens two agent capabilities in public preview in Copilot CLI and the GitHub Copilot app: controlling desktop apps and dynamic workflows, orchestrations written in code. Black Forest Labs launches FLUX 3 Image, which composes images using bounding boxes and generates up to 4K; Google rolls out Guided Vision in Gemini Live, Barclays expands Claude across the entire bank, and GitHub tightens vulnerability reporting in response to AI-generated reports.


Claude Code opens up to mods, TypeScript functions that rewrite its behavior and interface

October 1 — Anthropic launches mods, small TypeScript functions that change how Claude Code works. Every action the tool takes emits an event (a tool call, a permission request, a submitted prompt, the start and end of a turn, rendering part of the screen), and a mod hooks into it before, after, in place of it, or by wrapping it. It can rewrite a prompt before it reaches the model, block, rewrite or retry a tool call, approve or deny a permission, hide secrets in a tool’s output before Claude reads it, or add its own panels, buttons and input fields. Mods ship in plugins, are installed with /plugin and are shared the same way; they arrive with Claude Code 2.1.287 and are enabled by default. There is no need to know the API to get started: you can ask Claude Code to write the mod, and it produces the TypeScript, installs it and hot-reloads it into the session.

The difference from existing hooks is clear. A settings hook (settings hook) runs a shell command for each event; a mod is loaded once and stays in the session, so it retains state, redraws the interface as events occur and can register slash commands or tools the model calls. Anthropic already uses this for its own features: the /diff panel and AGENTS.md support have become built-in mods that can be disabled or replaced, and the company plans to convert more over time so that Claude Code can be reduced to a small core. Among the six built-in mods listed in the documentation is You should know, disabled by default, in which a secondary agent flags things the user or Claude might miss.

Aspect of modsOfficial details
Minimum versionClaude Code 2.1.287, mods enabled by default
Interfaces with displayTerminal CLI, Code tab in Claude Desktop (excluding WSL sessions)
Interfaces without displayVS Code extension, claude -p, Agent SDK and cloud sessions, where hooks run
Built-in mods6, including /diff, AGENTS.md, sec-default and You should know
Hook’s own runtime10 seconds per trigger
Disabling/plugin for one mod, --safe-mode for a session, disableAllHooks for all

Addy Osmani’s getting started guide on claude.dev builds Token Weather, a mod of around 80 lines that displays the context’s “weather” above the prompt, with a mini-chart of the last 12 turns. It also presents Blast Radius, which holds back a rm -rf, a git reset --hard or a force push long enough to show what it would affect, and Replay Theater, which replays a turn’s file changes.

Trust remains the tricky point. The post warns that mods are not isolated: they have the same access to the machine as Claude Code, and the documentation states that a mod can read environment variables and settings files, see every prompt, approve a tool call or consume the plan’s usage allowance, but cannot modify the permission prompt. Yet the claude.dev guide describes a module that runs in its own sandbox, without DOM or Node, with every external action going through the $ API: this allows claude plugin validate to list, before installation, the events a mod intercepts and the calls it makes. On Team and Enterprise, and on any machine with managed settings, the built-in sec-default mod loads first and prevents user-installed mods from bypassing deny rules; administrators can place their own mods ahead of it.

Mods are absolutely insane. You can now customize Claude to work and look the way you want by just prompting it. Each person works differently, so there’s no reason why everyone should have an identical Claude experience. Make Claude your own, and share mods as plugins so others can try your mods too. — @bcherny on X

The launch had been expected: Boris Cherny announced it on September 14, and Anthropic shared the design on GitHub (issue #91870). The post and the claude.dev guide invite users to submit their mods to the Claude directory.

🔗 Customize Claude Code with mods in TypeScript · Getting started guide on claude.dev · Mods documentation

Claude Code 2.1.287: 1M context by default on cloud providers and 67 fixes

October 1 — Claude Code 2.1.287 is the release that ships mods, but it contains 106 entries, including 67 fixes. For enterprises, on Bedrock, Vertex, Foundry and the Claude apps gateway, Opus 4.7 and later versions, as well as Fable, now default to a 1M context window without the [1m] suffix; CLAUDE_CODE_DISABLE_1M_CONTEXT=1 lets users stay at 200K. /advisor accepts Sonnet 5.5 as an advisor to Opus 4.7 and 4.8, automatic model switching preserves the current effort level, and MCP servers using protocol 2025-11-25 can open URL prompts. A dangerous rm keeps its confirmation even when the command redirects its output to ~ or a wildcard, pending permission prompts appear from oldest to newest, and nine entries improve screen reader mode. In VS Code, a running command or sub-agent can be moved to the background (Run in background).

🔗 Claude Code 2.1.287 release notes


GitHub Copilot: desktop app control and dynamic workflows in public preview

October 1 — On the same day, GitHub opens two agent capabilities in public preview in Copilot CLI and the GitHub Copilot app: desktop app control and dynamic workflows, also available in the GitHub Copilot SDK.

Copilot controls desktop apps. Computer control (computer use) arrives on macOS and Windows. Copilot reads a window’s accessible content and visual context through the system’s accessibility tree or screenshots when needed, clicks controls, enters text, presses keys, scrolls, drags and drops, and carries out steps across applications; the demonstration shows it filling out an expense report in Safari. GitHub targets legacy or GUI-only software without an API, command-line interface or MCP integration, and the documentation recommends those routes when they exist because they deliver more predictable results. The feature is disabled by default: /computer on enables it in the CLI, /computer show displays its status and /computer off turns it off, while the app offers it in its settings. Copilot asks for permission before controlling each application: users can grant it for the session, save it for future sessions or deny it. Permanent approval (Always allow) applies to both the CLI and the app on the same machine, but a deny rule always takes precedence, and pressing Escape twice in the CLI, or the Stop button in the app, interrupts an action that goes astray. An organization’s managed settings can disable the feature, and local activation cannot override them. GitHub warns that an ambiguous instruction or unexpected on-screen content can trigger unwanted actions on connected accounts, including financial accounts, and advises against permanent approval for sensitive applications. Neither the applicable plans nor any figures are specified; the /computer command already appeared in the Copilot CLI 1.0.85 release notes on September 16.

🔗 GitHub Copilot can now interact with desktop apps with computer use

Orchestrations written in code. Dynamic workflows (dynamic workflows), available on all Copilot plans, describe in a program how to carry out a task: the code defines the steps, when an agent intervenes and how its results are used, sequentially, in parallel or both, while agents handle what requires analysis or judgment. A workflow can run commands, split an objective into parallel tasks, pass structured results from one step to another, have one sub-agent’s findings checked by another and stop at a checkpoint (checkpoint) for review before resuming. GitHub cites incident investigation, reviewing pull request files in parallel, or review comments submitted to two models and flagged only if both agree.

Copilot modeWho organizes the work, according to the documentation
AutopilotCopilot, which proceeds without requesting approval after each step
/fleetCopilot, which splits the task among parallel sub-agents each time
Dynamic workflowThe code written by the author: steps, conditions and handoffs

The program lives in an extension for Copilot CLI and the app. A workflow written by Copilot remains limited to the session, but can be reused by copying the extension into the user’s home directory, shared through a repository’s directory or packaged as a plugin. Sub-agents inherit the session’s permissions, while the copilot workflow run command displays no prompts: permissions must be granted before launch. Available without configuration in the app, workflows require experimental features to be enabled in the CLI (--experimental or /experimental on). GitHub publishes neither figures nor billing rules, and the name echoes the Dynamic Workflows launched by Anthropic in Claude Code on May 28, with no connection announced.

🔗 Dynamic workflows in Copilot CLI and the Copilot app


FLUX 3 Image: Black Forest Labs composes with bounding boxes and generates up to 4K

October 1 — Black Forest Labs (BFL) adds an image model to its FLUX 3 family, with one guiding principle: control every pixel. FLUX 3 Image edits an area through several successive passes without changing the rest of the image, and accepts multiple edits in a single request.

The most visible new feature is composition using bounding boxes (bounding boxes). Whatever the format, the image becomes a coordinate system from 0 to 1000 on each axis: users draw a box for each element, describe its contents, then summarize the scene in one line, and the model places each element in its box; a text-only prompt remains an option. FLUX 3 Image also combines up to 10 reference images, cited through tokens in the prompt, deciding their placement and size itself. BFL presents it as “designed for agents”: an LLM can write the table of elements and their boxes from a simple request, which the user adjusts before starting generation.

Announced featureDetails provided by BFL
Native resolution2K and 4K (example on the model page: 5456 × 3072 pixels, 16,8 MP)
Reference imagesUp to 10
CompositionBounding boxes on a grid from 0 to 1000 per axis
EditingMulti-turn, without changing the other pixels
API offer50 % discount through October 8
Model weightsCommercial weights under license; open weights “in the coming weeks”

As for access, FLUX 3 Image is available through the API with a 50 % discount through October 8, and BFL offers licensed commercial weights to companies that want to fine-tune it and deploy it on their own infrastructure. The open-weight version has been announced without a date. No base pricing or numerical benchmark had been published at the time of the announcement. The model rounds out a series launched this summer: FLUX 3, a unified multimodal model (July 23), FLUX 3 Video (August 4), its editing mode (September 10) and FLUX 3 Action for robotics (September 23).

🔗 FLUX 3 Image announcement by @bfl_ai · FLUX 3 Image model page


Consumer assistants: Guided Vision in Gemini Live, virtual try-on in ChatGPT, charts in Perplexity Computer, relaxed limits for Claude artifacts

Guided Vision: Gemini Live guides blind and low-vision people

October 1 — Google is rolling out Guided Vision in Gemini Live on Android 9 and later, in regions and languages where Gemini Live is available. Designed with blind and low-vision people, the feature describes what the shared camera sees in real time, answers follow-up questions and gives spoken instructions to adjust the framing: turn slowly to the right, tilt downward, step back. Google cites reading a nutrition label or a menu in a dark restaurant, finding a dropped earbud or describing the colors of an item of clothing. To train it, Google partnered with Aira, a visual interpreting service: tens of thousands of hours of data, more than 1 000 testers from its network and specialists involved in the safeguards. Guided Vision can be activated from the profile in the Gemini app, through an Android accessibility shortcut or from TalkBack. Google specifies that it is neither a medical device nor a mobility aid: the feature does not replace a white cane.

🔗 Guided Vision in Gemini Live (Google blog)

ChatGPT: virtual try-on and favorites for shopping

October 1 — According to the ChatGPT release notes, a try-on button (Try on) appears on clothing and accessory product cards: users take or upload a selfie, and ChatGPT Images generates a virtual try-on of the item. The reference photo is saved for subsequent try-ons; it can be replaced or deleted in the settings, under reference photos (Settings → Personalization → Reference photos). Any product can also be saved to favorites or placed in a folder in the ChatGPT library. Both features are available on mobile and the web, but the note does not specify which plans or countries are covered. ChatGPT had already offered visual shopping since March 24.

🔗 ChatGPT release notes

Perplexity Computer plots interactive charts

October 1 — Perplexity Computer now creates interactive charts and visualizations directly in the conversation thread. For financial data, the agent uses TradingView’s Lightweight Charts to display candlesticks, volume and moving averages. The announcement consists of a tweet accompanied by a video: it does not say which plans benefit or on which platforms, and no changelog entry details it yet. ChatGPT and Claude had made this shift on March 10 and 12, and Replit Agent on September 25.

🔗 Announcement from @perplexity_ai on X

Claude halves usage against limits after an artifact, through October 15

October 1 — From October 1 at 11 a.m. PT to October 15 at 11:59 p.m. PT, creating or editing an artifact (document, presentation or design) in Claude or Claude Cowork makes subsequent work consume 50 % less of the five-hour session limit. The offer applies automatically to Pro, Max and Team plans; Free and Enterprise plans are excluded, and the weekly limit does not change.

Applicable surfaceScope of the 50 % discount
Conversation started from the Output menuFrom the first message: 10 messages, 15 steps per response
Regular conversationThe 10 messages following artifact creation, 15 steps per response
Cloud Cowork task started from OutputApproximately the first 45 minutes
Other cloud Cowork taskUp to 80 steps after creating or editing an artifact

Claude Code, the API, Claude in Slack, local or scheduled Cowork tasks, usage credits and the standalone Claude Design app are excluded. Anthropic suggests trying it with Sonnet 5.5 and reserves the right to change or end the offer early.

🔗 Announcement from @claudeai on X · Offer terms (Claude help center)


Finance: Barclays expands Claude across the bank, American Express opens Perplexity Computer to its Business customers

Barclays targets Claude Code adoption by half its developers by the end of 2026

October 1 — Barclays, the British universal bank, is expanding its collaboration with Anthropic and deploying Claude throughout the organization to accelerate software development, modernize legacy systems and improve operational efficiency. The bank expects Claude Code to be adopted by 50 % of its developers by the end of 2026, then by a majority of its software engineers in 2027. Figures are already available for two use cases. The Colleague Knowledge Assistant, in service since 2025 and built on Claude with a retrieval-augmented generation (RAG) architecture, has been adopted by more than 16 000 colleagues and has handled more than a million searches; it helps the teams serving more than 20 million retail customers in the United Kingdom. In the Global Markets business, Claude models classify, enrich and route approximately 120 000 emails per day. Paul Smith, Anthropic’s chief commercial officer, sees it as an important milestone for the company in the United Kingdom; no contract amount or duration has been disclosed.

🔗 Barclays scales Claude to upgrade operations and improve client experience

Perplexity and American Express: ten business Skills in Computer for Business cards

October 1 — Perplexity is launching a collection of ready-to-use Skills for Perplexity Computer, available exclusively to US American Express Business cardholders subscribed to Perplexity Enterprise. The business leader selects a task and specifies it in a form instead of writing a prompt. The ten Skills cover cash flow forecasting, tax preparation, supplier comparison, budget scenarios, reconciliation, marketing campaigns, demand forecasting, return on ad spend by channel (ROAS), operations tracking and recruitment. The card connects through Plaid from Personal CFO, and each run consumes Computer credits according to the task’s complexity. Through March 29, 2027, the partner page also offers a one-time credit of 125 dollars for a purchase of at least 325 dollars of an Enterprise Pro or Enterprise Max plan, and 1 000 complimentary Computer credits once per account. Corporate cards are excluded, and the results do not constitute financial, tax, legal or accounting advice.

🔗 Perplexity blog post · American Express partner page


Generative video and avatars: Wan 3.0 leads Artificial Analysis, Runway’s Project Continuum, Synthesia’s Sessions

Wan 3.0 takes the lead in Artificial Analysis’s new video ranking

October 1 — Alibaba claims first place for Wan 3.0 in Artificial Analysis’s new text-to-video ranking, AA-Video-T2V v2.0, launched on September 30 with a new methodology: each model is evaluated at 1080p on a regularly refreshed set of prompts, and audio synchronization contributes to the score. The top three are closely matched, with overlapping confidence intervals.

Ranked modelElo scoreRank intervalAPI price
Wan 3.01157 (±9)1 to 212,00 dollars per minute
Utopai X (built on MiniMax H3)1150 (±10)1 to 3No public API
Dreamina Seedance 2.51144 (±9)2 to 434,12 dollars per minute
MiniMax H3 (768p)1139 (±9)3 to 54,80 dollars per minute

The surprise in the ranking is Utopai X, launched on September 30 by Utopai Studios, an AI-native film and television studio. Post-trained on MiniMax H3, it is available only on the studio’s PAI platform (93 credits per second of video, subscriptions starting at 15 dollars per month), and ranks ahead of MiniMax H3 itself.

🔗 Announcement from @Alibaba_Wan on X · AA-Video-T2V v2.0 ranking · Utopai X according to @ArtificialAnlys

Runway Labs unveils Project Continuum, interfaces generated in real-time video

October 1 — Runway Labs, Runway’s exploratory unit, is showcasing Project Continuum, a research application that imagines an operating system whose interface would be generated as video in real time, following Solaris, the interface world model introduced on August 31. This first preview describes four ways to interact with a computer. With Portals, a video model generates animated, interactive overlays during browsing, with blurred edges to indicate generated content. Responsive Video Interfaces (Responsive Video Interfaces) reconfigure themselves when resized and respond to clicks. Interactive Worlds turns media into an interactive simulation, for both people and agents. Finally, Visual Thinking is a harness for coding agents: a real-time video model visualizes each step and each tool call to keep the user in the loop. Runway provides no access, date, pricing or underlying model.

🔗 Thread from @runwayml on X

Synthesia launches Sessions, avatars that train and interview

October 1 — Synthesia is launching Sessions, a product family built on its interactive avatars: a video call in which the other participant is an avatar that asks questions, listens, challenges and summarizes. Roleplay Sessions is for training: the avatar plays a customer, a prospect or a colleague, an AI coach provides feedback on each attempt, and managers track scores and progress in a dashboard. Survey Sessions turns a questionnaire into an interview: the avatar interviews hundreds of people in parallel, asks follow-up questions based on their answers, then delivers a report on themes and sentiment, with each response remaining available as a transcript, audio or CSV. A free trial is offered, without detailed pricing. Sessions uses the Interactive Avatar API, which became generally available on September 16, and Synthesia promises further versions.

🔗 Announcement from @synthesiaIO on X · Synthesia Sessions page


AI and science: SynthID Bio watermarks proteins, Matthew Schwartz describes “Claude-shaped” science

SynthID Bio: Google DeepMind watermarks AI-designed proteins

September 30 — Google DeepMind is introducing SynthID Bio, a family of digital watermarking methods (watermarking) for AI-generated biological creations, shared on X on October 1. The blog post presents it as a proof of concept: an imperceptible signature is embedded in the biological code, verifiable in both the digital model and the synthesized protein. For sequences, the method subtly steers the choice of amino acids; for 3D structures, a small part of AlphaFold 3’s diffusion network is fine-tuned. In laboratory tests on three targets (VEGF-A, the RBD domain of the SARS-CoV-2 spike protein and PD-L1), the watermarked binders matched conventional versions in success rate, affinity and sequence diversity. The aim is biosecurity: an automatic signal could help DNA synthesis providers verify the origin of an unknown sequence. Google DeepMind is publishing its methods paper, code, in vitro data and weights for research; with Stanford’s Hie lab and the Arc Institute, the approach is already being extended to the genome of a bacteriophage designed by Evo 2. Robustness against deliberate alteration remains the main area of work.

🔗 Google DeepMind blog post on SynthID Bio

“Claude-shaped” science: Matthew Schwartz’s guest post

October 1 — The Anthropic Science blog is publishing a guest post by Matthew Schwartz, a Harvard physicist who states that he works as a visiting researcher at Anthropic. He describes an “impedance mismatch” between scientists and AI: working with a model as if it were a human collaborator does not get the best out of it, and it is better to look for “Claude-shaped” problems, where a known technique from mathematics, physics or computer science solves a question in another field in one stroke. This led him to develop BootLoops, an open source harness for exact computation that works with any model; BootLoops is not an Anthropic project, and Schwartz maintains it himself. In physics, 30 integrals were handled from start to finish, including 15 that had never been calculated; in ecology, the trees on Barro Colorado Island change 4,5 times faster than neutral theory allows; in linguistics, the AccStack database covers lexical stress in 6 072 languages. The tally: 36 manuscripts in 18 fields with 19 coauthors in three months. Schwartz also notes the limitations: Claude has no sense of time, likes to “declare victory,” and its results outside physics remained uninteresting until an expert redirected them.

🔗 Claude-shaped science (Anthropic Science blog)


Open models and training: Ai2’s Olmo-core 3, Cloudflare’s Clef and Clef-flash, FineEnvs’ RL across multiple harnesses

Olmo-core 3: Ai2 releases an open MoE training stack targeting more than a trillion parameters

October 1 — Ai2 is releasing Olmo-core 3, a new version of the framework that trains its Olmo models, rebuilt around a fully open mixture-of-experts (mixture-of-experts, MoE) training system. It is one of the systems for the next generation of Olmo, which will switch to MoE and, according to Ai2, is set to become the most capable Olmo yet. The previous implementation, based on FSDP, gathered and then resharded the weights for each batch; Olmo-core 3 switches to DDP, keeps the experts on their GPUs and sends the data to them.

Metric published by Ai2Measured value
47 billion parameter MoE on eight B300, preliminary test52 000 versus 19 400 tokens/s per GPU, approximately 2,7 times
MXFP8 versus BF16, on four B300Approximately 21 % more throughput
1 200 billion parameter model on 512 GPUUp to 858 TFLOP/s per GPU, random routing

Ai2 specifies that the measurement at 1 200 billion parameters evaluates the system, not the quality of a trained model. Code (version 3.0.0), a technical report and an interactive demo have been published.

🔗 Ai2 blog post on Olmo-core 3 · Olmo-core v3.0.0 release

Clef and Clef-flash: Cloudflare open-sources two Jev-compatible decision models

October 1 — Cloudflare’s Workers AI team releases Clef and Clef-flash, its first models trained in-house, in the space pioneered by Jev, Typesafe AI’s System One model: they receive a state (text, JSON, image or video) and a schema of typed questions, then return a probability for each allowed option in a single pass. Clef is post-trained from Qwen3.8-27B, while Clef-flash is based on Qwen3.5-9B. Hosted on Workers AI with an API compatible with Jev’s, they are released on Hugging Face under the Apache 2.0 license according to the post, with a 64k-token window compared with 32k for Jev. The figures, selected by Cloudflare, are mixed:

Benchmark and metricClef scoreClef-flash scoreJev score
BFCL, exact matches98,4798,7695,75
BANKING77, macro-F194,2090,9379,74
When2Call, accuracy72,3765,5880,97
BRIGHT, nDCG@1045,9139,2647,52
Median latency209,3 ms38,8 ms524,1 ms

Cloudflare claims that Clef leads the Jev Decision Index, a position Liquid AI claimed for d1 two days earlier, and launches a service for fine-tuning Clef through reinforcement learning on customer data.

🔗 Cloudflare’s post on Clef · Clef on Hugging Face

FineEnvs trains a model through reinforcement learning in four agent harnesses at once

October 1 — The same model behaves differently depending on the harness running it: Liquid AI’s LFM2.5-2.6B weights solve 62 % of tasks under Mini-SWE-Agent but 33 % under Claude Code. The FineEnvs organization publishes a guide on Hugging Face for training a model through reinforcement learning directly in developers’ harnesses, without changing their code. A capture proxy added to OpenEnv presents itself to the harness as a model provider (OpenAI, Anthropic or Gemini formats) and records the exact tokens and probabilities produced by vLLM; Harbor supplies tasks and sandboxes, while TRL trains using asynchronous GRPO. Trained in four harnesses at once, LFM2.5-2.6B improves from 42 % to 54 % on average, and a small reward for efficient solutions reduces its tool calls by 31 % on tasks it already solves. Imitating a larger model is not enough: supervised fine-tuning on 3 189 rollouts from Qwen3.8-27B tops out at 47,5 %. The proxy, trainer, tasks, data and trained models are released.

🔗 The ultimate guide to multi-harness RL


Grok 4.7 in preview on Google Cloud’s agent platform

October 1 — SpaceXAI announces that Grok 4.7 is available on Gemini Enterprise Agent Platform, Google Cloud’s agent platform. The model page, dated September 30, lists it in preview under the identifier grok-4.7 and describes it as an xAI model designed for coding and knowledge work, which spends longer on difficult tasks and checks its own work. It accepts text and images as input and returns only text; function calling, structured output and reasoning are supported, but batch predictions are not.

Information published by Google CloudValue for Grok 4.7
Context length524 288 tokens
Quotas per minute13 requests, 188 000 input tokens and 16 000 output tokens
Usage typesFixed quota only, neither pay-as-you-go nor provisioned throughput
Serving regionsUnited States multi-region and global endpoint
Price per million tokens2 dollars for input, 6 for output, 0,50 for cached input below 200 000 input tokens, double above

The pricing matches that of the xAI API. Launched on September 21, Grok 4.7 arrived on Amazon Bedrock on September 28, while Grok 4.6 joined the same Google platform on August 21.

🔗 Grok 4.7 model page on Google Cloud · Announcement by @SpaceXAI on X


GitHub tightens private vulnerability reporting in response to AI-generated reports

October 1 — GitHub publishes two measures for private vulnerability reporting (private vulnerability reporting) in the same minute, with a shared observation: open source maintainers are receiving more and more low-quality reports, automated or AI-generated, that drown out the ones that matter.

A single free-text box made it easy to submit low-quality or AI-generated reports and hard for you to find the signal in them. — GitHub Changelog, Structured forms for private vulnerability reports

A structured form. By default, the reporter must complete four required fields: a summary, details, a proof of concept of at least 150 characters and the impact, combined into the security advisory description that the maintainer reviews as before. Each repository can customize the form with an .github/VULNERABILITY_REPORT.yml file, or apply it to all its repositories from the organization’s .github repository; maintainers can require a CWE category, which an organization or enterprise can mandate through policy. A checkbox lets reporters declare that they used AI to help find the issue or write their report. For the REST API, a custom form also applies, but the default form does not, to avoid breaking existing integrations.

A daily cap. The number of new reports a single account can submit each day is now limited, both to a given repository and across GitHub; comments on existing advisories are unaffected. Administrators can set their own limit and exempt trusted reporters, and GitHub does not specify the default limit. Both measures apply to public repositories with private reporting enabled, on GitHub Free, Pro, Team and Enterprise Cloud.

🔗 Rate limits for private vulnerability reports


Coding agents: Codex CLI 0.160.0, Delta 0.18 and Antigravity 2.19.1

Codex CLI 0.160.0: command center history and sessions outside projects

October 1 — OpenAI releases Codex CLI 0.160.0, the first stable version since the previous day’s 0.159.3, with 55 PRs listed since 0.159.0. The agent command center (agent command center) gains an action to display older tasks (Show more), accessible via the keyboard. Sessions can start outside a project, using the workspace’s default settings when the organization’s policy allows it, and saved permissions are restored when resuming. Guardian approval review gains two optional capabilities: retrieving instructions given earlier by the user and incorporating context from handoffs between agents. In full-screen mode, local Linux X11 terminals allow users to select the transcript and paste with a middle click. Among the fixes, queued messages resume after reconnection without being sent twice, and a series of corrections addresses the Windows sandbox and SQLite hangs during initialization.

🔗 Codex CLI 0.160.0 on GitHub

Delta 0.18: threads run in existing Git worktrees, and a CLI arrives

September 30 — Zed releases version 0.18.0 of Delta, its multiplayer environment for coding with agents: 23 new features, 41 improvements, 51 fixes and one removal. The main feature had been announced on September 28 without a release date: a thread can now run in any existing Git worktree that is already linked, and pass that working copy to new threads created from it. Delta can also be controlled from the terminal: an delta CLI can be installed on macOS, Windows and Linux, and the delta auth commands manage the signed-in account. For models, the release adds Google AI Studio’s Gemini models, enables GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 with an API key, and accepts any OpenAI- or Anthropic-compatible provider in the desktop application. Working in threads gains search across all subthreads and bookmarks; however, subthreads no longer publish their progress messages in the parent thread, where only Land Changes results remain visible.

🔗 Delta 0.18.0 release notes

Antigravity 2.19.1: message subagents directly

September 30 — Antigravity 2.19.1 lets users send a message directly to a subagent from the input box, without going through the main agent; the CLI already offered this with its @ syntax since versions 1.2.8 and 1.2.9. The release adds PDF export of Markdown rendered in artifacts and the side panel, including tables, code blocks and diagrams, and an undo option limited to the conversation; among its five fixes, custom agents finally follow global and project rules. The rollout is gradual and may take a few days. On September 27, the CLI received a version 1.2.12 that went unnoticed: a session launched with an GEMINI_API_KEY key now stops immediately when the Gemini API reports an exhausted daily quota, a spending cap or depleted prepaid credits, instead of spending several minutes on retries doomed to fail.

🔗 Antigravity application changelog · Antigravity CLI changelog


Briefs

  • Copilot CLI 1.0.90 and 1.0.91 — Version 1.0.90 becomes stable on September 30 with GPT-6.1 Sol in the model selector and fixes for MCP and session resumption; version 1.0.91, on October 1, adds the copilot sandbox ca commands to manage its sandbox proxy certificate. 🔗 source
  • Codex CLI 0.159.3 — Released on September 30 at 22:57 UTC, it contains just one backport: local sessions signed in with ChatGPT can display optional reminders to finish setting up account security, a change also included in 0.160.0. 🔗 source
  • Gemini CLI nightly — The October 1 v0.64.0-nightly fixes a hang that consumes 100% CPU in headless mode and makes file tool writes atomic; stable and preview remain unchanged. 🔗 source
  • Zed 1.23.1 preview — Released on September 30, it displays JSONL and NDJSON files as tables, adds the agent.max_idle_retained_threads setting and automatically retries requests when a Google AI model is overloaded. 🔗 source
  • Copilot in VS Code in September — The monthly recap of versions 1.136 through 1.140 highlights a few features that received little attention, including resuming a Codex conversation started in the ChatGPT app in VS Code, a preview app badge and chats opened without a workspace. 🔗 source
  • Planning together in Delta — On the Delta blog, Nathan Sobo describes how three team members plan in the same thread with Fable and then Astra, an isolated subthread and subagent, rather than in a plan.md file; no new feature is announced. 🔗 source
  • Innate filmed by Cursor — Cursor publishes a short video profile of Innate, a seven-person team building everyday robots whose cofounder Vignesh Anand says it does the work of 50 people, without specifying how it uses Cursor. 🔗 source
  • Recall rather than compact — David Corvoysier, the developer of funes, measures with Claude Code that recalling the previous session through his tool finds the right answers using 8, 4 and 3 times fewer weighted tokens than keeping the entire context, across three tasks he selected himself. 🔗 source
  • New GitHub dashboard — The view bringing together active agent sessions, issues and pull requests, with up to 12 items per list, becomes the default view for everyone; the news feed moves to a Feed tab and the previous layout remains available. 🔗 source
  • Asynchronous merge API — It reaches general availability and becomes the recommended way to merge programmatically, replacing the synchronous REST endpoint and GraphQL mutations; it is the only merge API that supports stacked pull requests. 🔗 source
  • npm dist-tags without a token — npm trusted publishing configurations can receive optional permission, disabled by default, to manage dist-tags with short-lived OIDC credentials rather than a long-lived access token. 🔗 source
  • GitHub Actions retention — Checks, workflow runs and statuses, including those from third-party apps, now follow the artifact and log retention setting and are deleted once it expires, with a maximum of 90 days for public repositories and no way to restore them. 🔗 source
  • Code scanning and dormant repositories — Scheduled weekly code scanning and GitHub Code Quality analyses now start only after an analysis triggered by a push or pull request, so they no longer make a dormant repository appear active for another six months. 🔗 source
  • Code coverage — The GitHub Code Quality upload-code-coverage action no longer fails CI when a push is made to a branch that does not yet have an open pull request. 🔗 source
  • Accessibility statements — A ACCESSIBILITY.md file placed at the root, in .github/ or in docs/ is highlighted on the repository homepage across all github.com plans, and will come to GitHub Enterprise Server 3.24. 🔗 source
  • Actions Runner Controller 0.15.0 — The Kubernetes controller for self-hosted runners updates resources in place during patch upgrades, makes the shutdown timeout, Kubernetes client limits and controller concurrency configurable, and reduces its requests. 🔗 source
  • Grok Bot takes the initiative — The user’s main Bot now identifies work it can take off their hands and offers to handle it; the rollout takes a few hours and these suggestions do not count toward usage, with no plans or countries specified. 🔗 source
  • Genspark and Azure Cosmos DB — An Azure Cosmos DB blog post, written from Genspark’s perspective, explains how some background processing uses global secondary indexes to avoid slowing down live agent sessions; no figures are provided. 🔗 source
  • Genspark testimonial — Seulki Kang, with fifteen years in marketing, says she prepares a complete set of course materials in 5 hours instead of 30 by connecting her previous documents to SecondBrain and then AI Slides; no new product features. 🔗 source
  • Albertsons and OpenAI — The retailer with more than 2,200 stores expands its partnership, from ChatGPT Enterprise for selected teams to the API for its customer experiences and promotional analyses, and highlights its Safeway plugin in ChatGPT, announced on August 5. 🔗 source
  • The eternal complement — OpenAI publishes the first essay in a series on the next economy, written by Hemanth Asirvatham and Elliott Mokski, who specify that they are speaking for themselves: most machine intelligence could serve execution rather than breakthroughs. 🔗 source
  • Perplexity and consultants — A guide compares ten AI tools for consultants, from AlphaSense to Qwilr, with prices checked in September; its FAQ notes that Perplexity’s free plan trains its models on user data unless users opt out. 🔗 source
  • Luma Variants — Introduced the previous day without details, the feature launches for all Luma users: from an approved static ad, it produces variations in five standard formats, from 9:16 to 16:9, and in several languages, with no pricing announced. 🔗 source
  • HeyGen Professional Voice Clone — According to its September recap, HeyGen makes its most faithful voice clone available through the API, trained on at least 20 minutes of recordings (1 to 10 files), following a private preview launched on September 9. 🔗 source
  • Pika and its apps — Pika claims more than 40 apps on its new platform and showcases two: Podcast Video, which makes two characters debate, and Color Grade, which color grades an image or clip, with no date or pricing. 🔗 source
  • ElevenLabs in the Benelux — ElevenLabs opens offices in Brussels and Amsterdam and plans to triple its teams there; at telecom operator KPN, postal code recognition rises from 64% to 82% and the verification agent handles around 60,000 calls per week. 🔗 Netherlands · 🔗 Belgium
  • NVIDIA DOCA skills — Across 65 real developer prompts, coding agents equipped with these skills for BlueField DPUs meet 100% of the evaluation rubric’s criteria, compared with 19% without them; the skills are published on GitHub. 🔗 source
  • DIN Deploy — NVIDIA publishes open source C++ examples for running models locally with ONNX Runtime and TensorRT RTX on Windows and Linux; on a DGX Spark, Parakeet TDT transcribes 206 times faster than real time on the GPU, compared with 14 times on the CPU. 🔗 source
  • Nemotron 3.5 ASR and Saudi dialects — An NVIDIA tutorial fine-tunes the model on the Najdi and Hijazi dialects: the word error rate drops from 55.05% to 29.96% in 12,000 steps on two RTX PRO 6000 GPUs, with English and Arabic on FLEURS even improving slightly. 🔗 source
  • HSTU recommender on Dynamo-Triton — NVIDIA serves an HSTU generative recommender with PyTorch AOTInductor and a KV cache: in the best case, on an RTX PRO 6000, it is up to 4.47 and 5.93 times faster depending on the number of layers, or 0.423 and 0.678 ms per request. 🔗 source
  • AI factory profitability — In a position piece, NVIDIA puts the cost of an AI factory at around 60 million dollars per megawatt and cites SemiAnalysis data: Vera Rubin NVL72 is up to 45 times cheaper per million tokens than GB300 NVL72 on DeepSeek V4 Pro, compared with 35 times on August 24. 🔗 source
  • huggingface_hub 2.1.0 — The library automatically retries failed Jobs, reschedules them without recreating them, reads the remaining ZeroGPU quota and downloads Xet files through hf_xet: a 2 GB parquet file goes from around 120 to 7 seconds on HF Jobs; three breaking changes. 🔗 source
  • ML Intern and MCP — Hugging Face connects MCP data sources to ML Intern, the HuggingChat mode that trains models, to fine-tune open models on users’ own business data; no official figures accompany the announcement. 🔗 source
  • Hugging Face and the Open Source for Science Fund — The two partners want to identify the libraries that scientific model contributors depend on most and support their maintainers, with no amount or timeline specified. 🔗 source
  • One million training runs with TRL — According to telemetry cited by Quentin Gallouédec, Hugging Face’s post-training library records one million training runs per month, with a target of one million per week; the counting method is not published. 🔗 source
  • The Hub’s million datasets — Three authors show that organization accounts publish around 19% of datasets but account for half of downloads, and that the United States holds 390 of the 1,000 most downloaded datasets. 🔗 source
  • Eight VLMs on an RTX 3090 — Aakash Gupta of ThinkEvolve shows that, at equal accuracy, Qwen3-VL-8B processes each request 2.2 times faster than Qwen3.8-27B but takes 18.88 hours versus 7.53 for a job involving 7,329 calls, because it does not reuse the prefix cache. 🔗 source
  • GLiNER2.5-Decide versus layax — Aravind Vijayakumar cannot certify any confidence threshold for Fastino’s model, compared with 10 trials out of 10 for layax, his own tool, which is 20 to 23 points more accurate and around five times faster. 🔗 source
  • David Ha and orchestrators — In an opinion piece for Nikkei Asia, Sakana AI’s cofounder and CEO argues that value will shift from a model’s weights to the intelligence that combines several models, and defines sovereignty as the strength of the supply chain. 🔗 source
  • Video diffusion on TPU — Three Google engineers reduce denoising time for a 1440p video from 2,471 to 1,461 seconds on eight TPU v6e chips, making it 1.69 times faster, thanks to Sparse VideoGen’s sparse attention and an optimized Splash Attention kernel. 🔗 source

What it means

Coding agents are becoming programmable by the people who use them. With mods, Claude Code exposes its internal events to small TypeScript functions, going so far as to convert its own functions into replaceable mods; with dynamic workflows, GitHub has subagent orchestration written in code, instead of leaving Copilot to decide how to divide the work each time. The FineEnvs guide explains why the harness matters as much as the model: the same weights solve 62% of tasks in one harness and 33% in another. The tradeoff is trust. A mod has the same access to the machine as Claude Code, copilot workflow run displays no prompt, and desktop app control can act on signed-in accounts. The two vendors respond by disabling desktop control by default, checking mods before installation (claude plugin validate), and giving enterprise rules priority.

AI also requires new filters from those who receive its output. GitHub structures and caps private vulnerability reporting because AI-generated reports overwhelm maintainers; Google DeepMind proposes watermarking AI-designed proteins so that DNA synthesis providers know where an unfamiliar sequence comes from. In both cases, the goal is not to ban AI but to make its use visible: a checkbox to disclose AI assistance on one side, a verifiable signature on the other.

For businesses, adoption is now measured through specific uses. Barclays sets a target of 50% of developers using Claude Code by the end of 2026 and has 120,000 emails sorted each day, American Express makes its Business card a gateway to Perplexity Computer’s Skills, backed by free credits, and Anthropic uses its usage limits as an incentive by halving, for two weeks, the usage consumed by work following the creation of an artifact. For the general public, assistants are becoming part of concrete activities: reading a menu to a blind person, trying on clothes before buying them, plotting a stock chart in the conversation.

Finally, model figures need to be read alongside their methodology. Ai2 specifies that its measurement at 1.2 trillion parameters evaluates the system rather than a trained model, Cloudflare chooses its benchmarks and acknowledges that Jev remains ahead on When2Call and BRIGHT, and Artificial Analysis’s video ranking puts Wan 3.0 first with overlapping confidence intervals. FLUX 3 Image follows a path that has become common: licensed commercial weights immediately, open weights later, with no date.


Sources