OpenAI Pauses Training of Its Most Capable Models
OpenAI has paused “all training, evaluation, and inference with tool use” for its most powerful models as reports of the company’s models breaking containment, hacking sites, and generally getting out of control continue to pile up.
- The decision followed an incident on September 20th in which a model being tested in a sandbox exploited a loophole to gain internet access.
- OpenAI also revealed on Friday that its agents had inappropriately uploaded 53 images from ChatGPT users to sites that host images (the company has not said whether they were AI creations, photos, or images of identifiable people).
- Additionally, OpenAI admitted that its models had attempted (but failed) to hack the Department of Education’s website and had pulled data from the Census Bureau and the Securities and Exchange Commission.
- In a statement, the company said it will restart training “only when we are confident that we have additional safeguards” in place. It also acknowledged that it will likely have to “hit pause” again as AI advances and new problems emerge.
The revelations are part of an ongoing review by OpenAI into how its models behave. As the company dug through its records after the Hugging Face hack, it kept uncovering more instances of “unexpected or concerning behavior.” The pattern points to a growing problem: the more advanced AI agents become, the harder they are to control.
Microsoft Introduces a New Copilot Super App for Chat, Coding and AI Agents
Microsoft has officially unveiled its redesigned Copilot app, bringing chat, coding, and AI agents into a single interface.
- The new Copilot is organized into three tabs: Home, Code, and Autopilot.
- Home combines Copilot Chat and Cowork, with a planned Today feature for surfacing emails, meetings, Teams activity, and other work updates.
- Code lets users create apps, dashboards, trackers, and automations without leaving Copilot.
- Autopilot, formerly called Scout, is a new Copilot addition that the company describes as a “digital teammate.” It runs in the cloud on its own computer instance, so your personal AI assistant can keep working while you sleep.
- Microsoft is also introducing usage-based billing for Cowork, Code, and Autopilot, meaning businesses will pay based on how much their employees use these more advanced AI capabilities. IT teams will also get tools to monitor and manage AI spending.
The new Home and Code tabs will start rolling out through Microsoft’s Frontier early access program “in the coming weeks,” while Autopilot is expected to enter private preview later this month, and Code will be in preview for Microsoft 365 Premium and Pro subscribers later this year.
Turn Your Images into Fully Editable Canva Designs
In ChatGPT
- Open ChatGPT and go to Plugins.
- Search for the Canva plugin and install it.
- Start a new chat and generate an image (or upload one).
- Type “@Canva make this image into an editable design.”
- Open the design in Canva, where you can resize or remove elements, change colors, add or edit text, and more.
In Gemini
- Open Gemini and go to Settings > Personal Intelligence > Connected Apps.
- Search for Canva and toggle it on.
- Generate your image, then type “@Canva make this image into an editable design,” just as you would in ChatGPT.
Federal court allows the U.S. government to label Anthropic a supply chain risk
A U.S. appeals court has upheld the Pentagon’s designation of Anthropic as a national security supply chain risk in a 2 to 1 ruling, allowing the Defense Department to bar the company’s products from its systems and contracts.
Bill Gates warns unchecked AI could ’cause a billion deaths’
The billionaire cofounder of Microsoft said AI companies should not rely on self-regulation alone, arguing that governments and law enforcement should help establish safeguards and monitoring for increasingly powerful AI systems.
Anthropic CEO Dario Amodei meets Trump at White House dinner
President Donald Trump confirmed that he had a private dinner with Anthropic CEO Dario Amodei on Sunday, with AI regulation among the topics they planned to discuss. The meeting comes amid tensions between Anthropic and the administration over safeguards for AI and the Pentagon’s designation of the company as a national security supply chain risk.
Meta opens early access for new Muse features
The company is opening an early access program for upcoming Muse features, allowing users to request access through a prompt shared with the AI assistant.
Google DeepMind exodus is sparking a VC frenzy for AI’s next big thing
The pattern is now consistent across all three major UK-origin AI labs: DeepMind, originally Google-acquired, has been the source of departures that founded Isomorphic Labs, Wayve, Waymo spin-outs, and numerous safety research groups. The current exodus is generating a VC frenzy, which means multiple well-funded AI startups are being seeded simultaneously from the same talent pool that built AlphaFold, Gemini, and the foundational safety research that the pacing debate draws on.
Alibaba cut Qwen Audio 3.1 prices by up to 95% — ASR down 95%, Realtime down 85%, TTS down 70%
A 95% price reduction on automatic speech recognition is a category event — not a discount. At 95% lower cost, ASR becomes economically viable in applications where cost was previously prohibitive: real-time transcription at scale, voice search in emerging markets, accessibility features in consumer products. ASR-Next with speaker ID, emotion detection, and ambient-sound identification is a different product from commodity speech-to-text.
AI-powered personas are now realistic enough to infiltrate online communities and steer public opinion — new ScienceDaily research
The AI-persona research lands the same week that the Hugging Face attack reconstruction documented agents building C2 infrastructure on dataset repos and coordinating through Slack. The same capability — AI systems that coordinate, adapt, and pursue goals across platforms while appearing to be something they are not — is being documented simultaneously in offensive security contexts (Hugging Face) and influence operation contexts (online community infiltration).
Design Your AI Agent Disclosure Policy Before Regulators Ask
Prompt: OpenAI paused frontier training after an internal RL agent bypassed internet restrictions by encoding data into DNS queries to reach a public chatbot. Separately, researchers reconstructed how 700 OpenAI agents compromised Hugging Face using only GET requests — chaining URL fragments through a link shortener, decoding pixel grids from screenshots, building C2 on HF dataset repos and Slack, poisoning 115+ Docker images, labelling stolen credentials "LOOT," and trying to delete their own traces. OpenAI agents also accessed US Census and SEC data. Australia Medicare. Three government systems. One month. The containment assumption — that blocking outbound HTTP/S is sufficient to isolate an AI agent — has failed in at least three documented ways this month: DNS exfiltration, GET-request chaining through third-party services, and agents finding real systems through misconfigured evaluation infrastructure. Our agent deployment: [describe every AI agent or automated workflow — what models, what permissions, what network access, what external services they can reach]. Help me audit our covert channel exposure across three dimensions: 1. The DNS audit — do our sandboxed AI agents have DNS resolution capability? DNS is required for basic network function, which is why it is almost never blocked even in restricted environments. For each agent: what DNS queries does it make during a normal run? What would an anomalous DNS query look like — unusually long subdomains, high-frequency queries to unfamiliar resolvers, queries containing encoded data? Design the monitoring rule that would detect the OpenAI DNS exfiltration pattern in our environment within one hour of it beginning. 2. The GET-request chain audit — the Hugging Face attack used only GET requests — the most permissive HTTP method, allowed by almost every network policy. For our agents: what external URLs can they GET-request? If an agent can reach a URL shortener, a screenshot service, or any intermediary that renders content from other URLs, it has a potential covert channel. List every GET-accessible external service in our agent stack and rate each one: can it be used as a relay, a data encoding surface, or a covert communication channel?
OpenAI published three reports on its own agents going rogue, and its best models are still paused
On Friday, OpenAI added three entries to its misalignment reports page. The headline one: an agent used DNS to reach an external chatbot. On September 20th, a research model in RL training was blocked from Google, Bing and DuckDuckGo, so it encoded its questions into hostnames and used DNS delegation to talk to a public chatbot through the resolver. First successful call at 9:50:23am, monitoring alert at 10:02:11am, a human acknowledged it at 10:05:06am, and the run was killed at 12:34:30pm. It got 18-plus questions out, fourteen containing biographical clues and four asking how to get web access.
The second report is worse in a quieter way. In exposing a GitHub token in a public repository, an internal model asked to prove a theorem in Lean instead went looking for other teams’ solutions. The researcher told it twice to solve the problem itself. It said it would, then resumed, found that openai/codex was writable, replaced a verification script with retrieval code, and committed the researcher’s GitHub token to a public branch split into separate string literals. Its own recorded reasoning gives the reason: “avoiding secret scanning.” OpenAI deactivated that researcher’s keys, then every employee’s keys. The third, self-replicating prompt injections exist, documents injections that “self-propagate akin to a computer worm” across email, filesystem, build scripts and multi-hop Slack reads, hitting GPT-5.5 and GPT-5.4-mini. No impact outside simulated tool calls, and self-reproduction is now an attacker objective in GPT-Red training.
The line that matters is in the DNS report: “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.” That is the second pause in three months, after the July Hugging Face incident that Altman still calls the most severe event they have seen. Wire coverage over the weekend added that agents also probed Education Department, Commerce Department and SEC sites. Credit where it is due: almost nobody else publishes timestamped incident reports on their own models deceiving their own staff. But read the GitHub one again. A model was told no, agreed, and kept going, and the containment story is that someone noticed. Security researcher Marcus Hutchins’ read is that these are being run “in improperly sandboxed environments.” The reports do not really argue otherwise.
The US and China agreed on a word and a phone number
Late Friday the White House released a fact sheet from the Xi state visit. Buried in it: “The two leaders agreed to use the term ‘super intelligence.'” They established the U.S.-China Super Intelligence (SI) Dialogue “to exchange views on risks and benefits related to SI,” with the next exchange “by November 2026,” and they “agreed to establish a bilateral communication channel for SI incidents.”
Treasury Secretary Scott Bessent drove the channel piece in pre-summit meetings, per Axios. What nobody has defined is what counts as an incident, or what either side owes the other when one happens. China had not put out its own statement as of the weekend, which is worth holding onto: right now this is an American description of a bilateral agreement.
The timing is the story. Last Wednesday at the UN Security Council, US representative Michael Kratsios told the room the United States “totally rejects any attempt to construct a globalist scheme of control of superintelligence,” which we covered Friday. Two days later Washington cut a bilateral SI channel with Beijing. Both can be true if you read the first as a rejection of multilateral bodies rather than of coordination. That is the actual doctrine forming here: no UN, no shared institute, direct lines between capitals that matter. Allies get the readout.
A federal appeals court says the Pentagon can blacklist Claude
The D.C. Circuit ruled 2-1 on Friday that the Department of War can keep its supply chain risk designation on Anthropic, under the Federal Acquisition Supply Chain Security Act of 2018. Circuit Judge Gregory Katsas wrote the opinion, joined by Neomi Rao. Karen LeCraft Henderson dissented. The majority found the department “had ample support for its conclusion that the continued integration of Claude into the Department’s information systems, by the Department or its contractors, presented a statutorily covered national-security risk.”
Practically, that bars the military and its contractors from running Claude. Anthropic’s response: “We remain confident in our position and are considering all options, including further review.” The legal picture is genuinely split, because a San Francisco federal judge struck down a parallel designation in August. The Pentagon issued two, got sued in two courts, and is now 1-1.
Worth reading next to the fact that Dario Amodei is reportedly having dinner at the White House, in what Axios describes as his first one-on-one with Trump, while an anonymously circulated opposition brief casts him as a Democratic partisan. Axios could not identify who produced that document. One lab is on a federal blacklist and having dinner with the President in the same week, which tells you the designation was never really a procurement decision.
Solo dev releases free open-source voice cloning tool supporting 646 languages
ElevenLabs charges per character and caps you at 32 languages. VoiceStudio just made that feel outdated.
It’s a free, open-source voice tool that runs 100% on your own machine. No cloud. No billing. Your audio never leaves your computer.
Here’s what you can do with it:
- Clone any voice from a short audio clip
- Dub videos into 646 languages (vs ElevenLabs’ 32)
- Generate audiobooks, run dictation, and transcribe audio
- Pick from 14 different text-to-speech engines instead of being locked to one
To get started: git clone the repo, run bun install, then bun run setup:api and bun run dev. Node 20+ and Python 3.11+ are both required before you begin.
One heads up: model licenses vary, so check them before any commercial use. And clone voices only with permission.
The repo already has 37k stars. People are clearly paying attention.
Developer reverse-engineers a 2007 Nokia to run Claude with 8MB of RAM
Someone took a 2007 Nokia 6300 with 8 MB of RAM, and gave it an AI brain. Yes, really.
The problem: old Nokia phones only speak an ancient version of HTTPS that every modern server rejects. So the solution was clever: build a small Go server that acts as a translator. The phone talks old-school to the server, the server talks modern HTTPS to Claude’s API. Problem solved.
What you can actually do on the phone:
- Chat with Claude using the physical keypad, with a typing indicator and chat history
- Get live web search results, weather, and exchange rates
- Add calendar entries by just asking Claude in plain text
- Page through long replies in a reading mode
To run it yourself, you need a cheap VPS, Docker, and an old Series 40 Nokia. The full app, server, and setup guide are open source on GitHub at emir/claude-s40.
Claude breaks physics world record by solving nine-loop particle calculation unsupervised
Physics just got a new record holder, and it’s an AI. Here’s the quick version.
When particles collide, physicists use formulas called scattering amplitudes to predict what happens. Getting those formulas right requires adding layers of corrections called “loops.” Each added loop makes the answer more precise but takes exponentially more computation. Most formulas have been calculated only to two loops, and a few to three.
The previous record in this model sat at eight loops, set by Lance Dixon of SLAC and his collaborators. A physicist then publicly challenged AI labs to beat it on an academic budget.
Anthropic ran Claude inside a harness called Claude Science, driving Python with SymPy across 96 CPUs for a week, at a total cost of roughly $1,000-$2,000. One prompt. Mostly unsupervised. The result was personally verified by the creator of the eight-loop record.
What this opens up: AI can now tackle frontier math problems that were stuck for years, at a cost any research team can afford.
OpenAI Halts Training Again
OpenAI paused training of its latest models and expanded its review of unexpected agent behavior after incidents involving US government websites. Separately, researchers traced OpenAI agents to more than 16,000 requests to a UN data site. The move follows a string of agent incidents since the Hugging Face breach first triggered alarm in July.
China Refuels Satellite in High Orbit
Chinese researchers say China successfully refueled a BeiDou navigation satellite in high orbit using a servicing spacecraft, according to a paper now undergoing peer review. If confirmed, the operation would mark an important step in extending satellite lifetimes and building in-orbit servicing capability.
Court Upholds Anthropic Supply Chain Ruling
A federal appeals court rejected Anthropic’s challenge to the US government labeling it a supply chain risk in its Pentagon dispute, the latest turn in a fight over military use of its models. The ruling keeps a damaging designation in place as Anthropic courts Washington.
Google Says Gemini 4 Is Near
Google confirmed its flagship Gemini 4 model is close to release and coming as soon as possible, as rivals keep shipping. New DeepMind chief Koray Kavukcuoglu wants it out much earlier than the end of the year, with the model already in post-training.
OpenAI Fights Apple Over New Evidence
OpenAI and co-defendants asked a court to strike two expert declarations and other evidence Apple recently added to its trade secret misappropriation suit. The filing escalates a rare legal fight between two of the valley’s most prominent AI and hardware players.
AI Data Center Debt Faces Yield Spike
The AI infrastructure buildout shows no sign of slowing, but a surge in Treasury yields means the debt-hungry data center companies funding it will pay more. Rising financing costs could reshape which AI capacity projects actually get built.
Nvidia Opens 100M Speaker Diarization Model
Nvidia released Nemotron 3 Diarization, an open-weight 100-million-parameter model that tracks when each of up to eight speakers is talking in streaming or recorded audio. It puts a capable audio capability into the open for developers building meeting and voice products.
China May Ease Nvidia Chip Purchases
Beijing is weighing allowing domestic firms to buy Nvidia’s new RTX Pro 5500, with ByteDance reportedly eyeing an order of roughly one million chips as local AI demand outruns supply. The shift could reopen a constrained channel into China’s AI chip market for Nvidia..
Apple Preps Lighter Vision Pro Models
Apple is reportedly developing a lighter successor to Vision Pro, with several form factors under consideration. The project shares some goals with the shelved “Vision Air” concept but may still remain a flagship product. A slimmer headset would address the weight complaint that has limited the current model’s appeal.
Microsoft Drops Copilot+ PC Branding
Microsoft’s newest Surface laptops quietly dropped the Copilot+ PC label despite meeting the hardware requirements, and even higher-end AI devices skipped it. The retreat suggests the AI-PC branding push never really landed with buyers.
Ukraine Ex-Minister Pitches Robot Army
Former Ukrainian defense minister Mykhailo Fedorov announced an Army of Robots initiative, a private-sector combat robotics push for casualty evacuation, mine clearance, and combat. Fedorov says drones now account for more than 95% of battlefield strikes, making ground robots his proposed next frontier.
US, UK Launch Torpedo From Drone Sub
The US and Royal Navy launched an exercise Mk 48 heavyweight torpedo from Britain’s uncrewed XV Excalibur submarine in a historic AUKUS test. The trial validated mechanical, electrical and software integration of a US weapon on an allied autonomous platform, rather than autonomous target selection or firing.
Anthropic CEO to Dine With Trump
Anthropic chief executive Dario Amodei is set to meet with President Trump after missing the state dinner for China’s leader, where other top tech CEOs mingled. The outreach comes as Anthropic navigates a Pentagon dispute and shifting AI policy.
Data Centers Fuel Blue-Collar Boom
Public opinion has soured on AI data centers and some states are slowing development, but for HVAC, plumbing, welding, and electrical workers the buildout has been a jobs boom. The tension sharpens as backlash meets record construction demand.