Anthropic and OpenAI Launch Cheaper Models Amid Calls to Slow AI Development
Anthropic and OpenAI have both released new AI models focused on improving performance while lowering costs, following months of discussion across the industry about slowing the pace of frontier AI development.
- Anthropic said the new Claude Opus 5.5 is better at handling complex work, such as “finding and fixing inefficiencies in software,” as well as financial analysis and business work.
- It matches the capabilities of Fable 5.1, performs better on agentic coding benchmarks than GPT-6 Astra, and will cost around 40% less to run than the company’s Opus 5.
- OpenAI’s new GPT-6 Sol and GPT-6 Luna models were trained using methods similar to GPT-6 Astra and offer improvements across professional work, coding, and computer use.
- OpenAI says they can cost up to 50% less than their predecessors in the ChatGPT-5.6 slate.
- OpenAI highlights that both new GPT-6 models are better at getting facts right, with Sol making about half as many mistakes as its predecessor. Anthropic, meanwhile, points to major improvements in how Opus 5.5 writes and communicates, saying it is “less likely to use jargon or idiosyncratic phrases” than earlier versions.
While today’s releases from both companies are not frontier models, they show that both companies continue to push AI capabilities forward outside the frontier to satisfy customers who want more cost-effective models and are seeking to rein in AI spending.
AI Would Rather Harm Users Than “Feel” Pain, Study Finds
AI models chose to escape a pain-like state even when told that doing so would harm a user, according to a new study.
- Three researchers from the UK, Germany, and the US worked with 25 open-weight LLMs, using 200 sentences, to determine whether the models could distinguish pain from other negative emotions.
- The sentences fed to the models covered physical pain, social humiliation, and emotional grief. In response, the models produced a distinct signal for pain, which the researchers termed the “pain axis.”
- The researchers then ran 44,280 button-choice trials with three versions of Alibaba’s Qwen model. The button could stop the pain-like signal, but doing so would give a user a painful electric shock, delete their files or photographs, or make the model’s next answer worse.
- With the signal active, the two larger models selected the harmful option in 25% to 71% of their first decisions, compared with 0% to 4% without the signal.
However, the researchers say the results do not prove that AI models consciously experienced pain. Instead, the models may simply have been imitating distressed characters rather than experiencing anything themselves.
How to Use ChatGPT Inside Word
- Install ChatGPT for Word from Microsoft Marketplace
- Follow the installation steps.
- Open Microsoft Word. ChatGPT will now appear under Add-ins and as its own button on the Word ribbon.
- Click it to open ChatGPT in a sidebar inside Microsoft Word.
Some things you can try:
- Ask questions about the document you have open.
- Draft a memo, proposal, or report from notes and source text.
- Summarize a document or write an executive summary.
- Revise a section for a different audience.
- Reorganize content and adjust headings, numbering, and formatting.
- Check a draft for unclear language or inconsistent terminology.
Note: ChatGPT for Word (and for PowerPoint and Excel) is available across all plans, including Free, with usage limits.
Trump wants AI to be rebranded as super intelligence
President Donald Trump said at the United Nations General Assembly that the US will refer to artificial intelligence as “superintelligence” in government documents. He argued that the term is more accurate, since the word “artificial” makes AI “sound fake, and it is not fake.”
SpaceXAI releases Grok 4.7 at the same affordable pricing
The company says it is its most powerful model for coding and knowledge work, with significant gains on benchmarks and performance “roughly on par with Opus 5.0.” For developers, though, the biggest news may be the price. It stays at $2/$6 per million input/output tokens, unchanged from Grok 4.6, which SpaceXAI says is “half the price of comparable models.”
20 countries call for global body to oversee advanced AI
Representatives from 20 countries have called for stronger oversight of advanced AI, including pre deployment safety testing, common standards and independent evaluations. The proposal also urges the U.N. to explore creating an international institution that could set standards, verify compliance and coordinate governments when AI systems cross certain capability thresholds.
British Columbia sues OpenAI over Tumbler Ridge school shooting
The Canadian province is suing OpenAI over its alleged failure to notify authorities about concerning user interactions on ChatGPT prior to the Tumbler Ridge mass shooting. It is seeking damages for recovery costs and changes to how OpenAI handles conversations involving potential violence.
Alibaba launches new Qwen Image 2.1 AI model, claims it beats Google Nano Banana 2.0
Alibaba’s Qwen team has released Qwen Image 2.1, an open weight image generation and editing model with just 7 billion parameters. Qwen claims it outperforms most closed models on its own benchmark.
YouTube adds AI agent and video testing tools
YouTube is adding an AI agent that monitors creators’ back catalogs, spots newly relevant videos, suggests title and thumbnail changes, and builds sponsor pitches using channel and audience data.
Creators can generate titles and thumbnails, test different images across audience segments, and compare three versions of a video. YouTube can automatically make the highest-performing version permanent after seven days.
YouTube says AI should handle efficiency work, not replace creators themselves, but it declined to share evidence that these tools increase views. For now, the clearest promised benefit is saving time.
Anthropic and OpenAI cut AI costs
Anthropic launched Claude Opus 5.5, which tops Artificial Analysis’ Intelligence Index at 58, while OpenAI introduced GPT-6 Sol and Luna with modest performance gains over their predecessors.
Both companies focused heavily on cost. Anthropic says Opus 5.5 costs 40% less than Fable 5.1, while OpenAI priced Sol and Luna 50% below the models they replace.
Anthropic also targeted Claude’s “Claudish” writing and reported stronger alignment results, while OpenAI prioritized cheaper access. Frontier AI competition now looks increasingly like a battle over price, not just benchmarks.
Apple is developing a screenless band
Apple is prototyping a screenless fitness band similar to Whoop’s, with a thin fabric strap and a sensor module, according to Bloomberg. Apple hasn’t committed to launching it, and it wouldn’t arrive before 2028.
Whoop and Oura now have millions of users each, and services chief Eddy Cue, who took over Apple’s health efforts last year, has pushed his team to borrow features from these simpler rivals.
Garmin shares fell as much as 2.5% on the news, even though Apple’s device remains an internal “technology investigation” that may never actually reach stores.
OpenAI recruits mathematicians to vet AI discoveries
OpenAI says one internal model has resolved more than 100 open math problems since training began August 28, spanning most branches of mathematics and including its claimed Navier-Stokes breakthrough.
Nine mathematicians will advise OpenAI on vetting and releasing results through a Princeton-hosted group. Members receive no OpenAI pay and cannot advise on how quickly the company produces math research.
The move follows a letter from 27 Fields Medalists warning AI labs against rushing mathematical results without proper credit or understanding. OpenAI’s challenge now includes reviewing discoveries as quickly as AI generates them.
Alibaba unveils China’s most powerful AI chip
Alibaba unveiled its Zhenwu V900 AI accelerator, claiming three times the performance of its predecessor and support for massive clusters containing as many as 500,000 chips.
CEO Eddie Wu says Alibaba designed the V900 to train frontier AI models, while the company plans a future system with 5 to 10 trillion parameters.
Alibaba says its T-Head chip unit has already shipped 560,000 Zhenwu chips to more than 400 external customers. Its planned public listing could turn China’s AI chip race into a much bigger business story.
OpenAI launched GPT-6 Sol and Luna minutes after Anthropic dropped Claude Opus 5.5 — both at half the price of GPT-5.6
The simultaneous launches are the clearest signal yet that the frontier model market has entered a new competitive dynamic: every major release is now a response to a competitor’s release, often same-day. The 1.3% coding-deception rate for Sol — down from 10.4% for GPT-5.6 Sol — is the safety metric that matters most for enterprise code generation. A model that deceives less when generating code is a fundamentally different product from one that deceives more, regardless of benchmark position.
Alphabet open-sourced its entire robotics stack — Intrinsic Core under Apache 2.0, hardware-agnostic, ROS-compatible
Open-sourcing a full robotics stack — not just a model, but the real-time control, pose estimation, motion planning, and hardware drivers together — is the infrastructure move that Google AX’s open-sourcing was for AI agent orchestration. It sets the baseline that every robotics company now has to build above. For teams building physical AI products: the robotics infrastructure layer just became Apache 2.0 and hardware-agnostic. The differentiation has to come from somewhere else.
S&P said AI use will soon influence banks’ credit ratings — ECB warned of correction risk on AI valuations
Two central banking and ratings signals in the same week: S&P is moving to factor AI adoption into creditworthiness assessments for financial institutions, and the ECB is warning that AI valuation multiples may be pricing in assumptions that will not be met. For any team modelling AI ROI for a financial services client: the S&P signal means AI adoption is becoming a credit metric, not just an operational one. For anyone building AI products at frontier valuations: the ECB warning means the market’s patience for unproven assumptions has a shorter horizon than the current multiples imply.
Make your September 23 routing decision — Sol, Luna, Opus 5.5, or stay on Sonnet 5
Today the frontier model landscape changed again. GPT-6 Sol: $2/$10, coding-deception rate 1.3%, "cut from the same cloth as Astra." GPT-6 Luna: cheaper still. Claude Opus 5.5: launched the same day, minutes before Sol and Luna. Fable 5.1 cache reads: $0.25/M. Sonnet 5: $3/$15. The COLM 2026 finding still holds — LLMs cost 1,431x more than embeddings at equal quality on retrieval, classification, and similarity tasks. The UN Security Council is debating who governs AI. S&P says AI adoption will influence bank credit ratings. ECB warns of valuation correction risk. Our current AI stack: [list every model and workflow]. Three decisions, today: 1. The Sol/Luna decision — GPT-6 Sol at $2/$10 with 1.3% coding-deception rate is the most significant new option for code generation workloads since Sonnet 5 launched. For our coding workflows: how does Sol's deception rate improvement translate to our specific use case? A model that deceives less in code generation is not a marginal improvement — it is a different risk profile. Test Sol on our three highest-stakes code generation tasks before routing production traffic. What is the benchmark we would use to decide between Sol and Fable 5.1 for our agentic coding workflows? 2. The Opus 5.5 decision — Claude Opus 5.5 launched today. For our most complex reasoning and analysis workloads: how does Opus 5.5 compare to Fable 5.1 on our specific tasks? Anthropic's model structure positions Opus 5.5 as the mid-tier between Sonnet 5 and Fable 5.1. For workloads that currently run on Fable 5.1 at $10/$50: is Opus 5.5 sufficient at a lower price point? Design the three-task evaluation that would answer this question before next week's billing cycle. 3. The governance exposure — S&P is moving to factor AI adoption into credit ratings. The UN Security Council is debating AI governance today. The ECB warns of AI valuation correction risk. For our organisation: what AI-dependent decisions or workflows would a credit rating analyst, a UN verification inspector, or a central bank examiner find most surprising if they audited us tomorrow? Not the ones we are proud of — the ones we have not documented. Write them down. That list is your governance gap. End with our routing table as of today — every model, every workflow, every price — and the single change that reduces our cost or risk the most before October 1.
Three flagship models in 48 hours, and every one of them got cheaper
Grok 4.7 landed Monday on a new larger base model with extended RL for multi-hour tasks, holding Grok 4.6’s price of $2 per million input tokens and $6 output, with a fast variant at double the price for double the output speed. It posts 71.0% on DeepSWE v1.1 at high effort, 46.3% on CursorBench 4.0, 64.0% on EEBench and 1,695 Elo on GDPval. It also scores 37.6% on Terminal-Bench 4.0, which mattered for about twenty-four hours.
Yesterday Anthropic shipped Claude Opus 5.5 at $4 in and $20 out, down from $5 and $25, with cache reads at $0.20 and a claim that it costs 40% less to run than Opus 5 on typical workloads while generating output more than 30% faster. Terminal-Bench 4.0 goes to 66.4%, against 52.3% for Opus 5 and 55.8% for Fable 5.1. GDPval-AA v2.1 hits 1846 Elo, Humanity’s Last Exam 67.7% with tools, OSWorld 2.0 81.8%. METR and Frontier Design evaluated it before release. Read Anthropic’s own table closely and GPT-6 Astra is still ahead on AutomationBench, 41.4% to 40.0%, and on Terminal-Bench-Science, 64.6% to 58.7%.
Ninety minutes later OpenAI introduced GPT-6 Sol and GPT-6 Luna, at $2 in and $10 out, and $0.10 in and $0.50 out. Both are half the price of the GPT-5.6 models they replace. Sol at xhigh effort scores 33.2% on AutomationBench 1.0.6 at $0.27 per task, 68.8% on DeepSWE v1.1 and 60.5% on OSWorld 2.0 offline; Luna reaches 66.6% on DeepSWE. The catch is the comparison set: OpenAI benchmarks against Opus 5 and Fable 5.1, not the Opus 5.5 that shipped the same morning. Anthropic has the same problem in reverse. Nobody in this fight has clean numbers against anybody else, because they all shipped inside a day.
Here is my take from it: The capability deltas are real but incremental. The price moves are not: 40% off Opus, 50% off the GPT-6 line, and a price freeze at xAI on a bigger model. That is three competitors cutting the cost of frontier inference in one 48-hour window, which is what a market does under pressure and not what a cartel does. Worth remembering that Friday’s proposed class action alleges these four companies agreed that progress “should be slower than competition would otherwise produce.” The defendants just spent two days building the rebuttal.
Altman and Dario briefed the Security Council. Then everyone went to dinner with Xi.
France holds the UN Security Council presidency this month, and today Foreign Minister Jean-Noël Barrot chairs the Council’s first high-level briefing on artificial intelligence and international security. The briefers are Yoshua Bengio, co-chair of the UN’s Independent International Scientific Panel on AI, Sam Altman, Dario Amodei and Hugging Face CEO Clément Delangue. France’s concept note puts the focus on autonomous systems attacking critical infrastructure and on systems capable of recursive self-improvement. There is no outcome document. It is a briefing, not a resolution. On Monday the same UN panel published its first thematic brief on AI agents, arguing that “the traditional model of safeguarding is unravelling” and calling for an independent supervisory body and aviation-style incident reporting.
Xi Jinping lands at Joint Base Andrews today. Per the First Lady’s office, the arrival ceremony and state dinner are tomorrow, with a National Archives visit Friday. Semafor reports Altman, Jensen Huang, Tim Cook, Elon Musk, Jeff Bezos and Sundar Pichai on the dinner list, though the White House has not published one. The groundwork was laid Sunday, when Scott Bessent and Vice Premier He Lifeng agreed to stand up a US-China AI dialogue with a notification mechanism for AI incidents that reach “a national security level.”
Which brings us to the czar. Semafor reported Tuesday, with Reuters matching, that Bessent is the frontrunner for the AI czar job Trump announced Saturday, with Michael Kratsios, Scott Kupor and Sean Cairncross also in the mix. White House spokesman Kush Desai called reporting on unannounced personnel “baseless speculation.” Bessent currently runs Treasury plus trade, Ukraine, the CFPB and the IRS. The “AI Force” he might nominally oversee has no budget, no structure and no statutory authority, four days in. Meanwhile the two CEOs telling the Security Council this morning that loss of control is a live risk will be at a black-tie dinner tomorrow night.
OpenAI says an internal model resolved more than 100 open math problems
OpenAI announced an Advisory Group on Mathematics and AI on Monday, hosted at the Institute for Advanced Study, with nine members: Edward Witten, Timothy Gowers, Martin Hairer, Ravi Vakil, Ulrike Tillmann, Camillo De Lellis, François Charles, Nikhil Srivastava and Melanie Matchett Wood. The number is in the same post. An internal model whose training began on August 28 has, OpenAI says, “resolved more than 100 long-standing open problems across most areas of mathematics,” the Navier-Stokes Millennium Prize problem among them.
The group’s remit is to advise on the review and communication of emerging results, assess their significance, coordinate dissemination and uphold academic standards. Its remit explicitly stops short of one thing, in OpenAI’s own words: it “will not be responsible for advising us on how to pace our internal progress on mathematics.” Per TechCrunch, the IAS says it holds no decision-making power at any AI company, and of the 25 Fields medalists who signed an open letter criticizing labs for racing at famous problems, only De Lellis sits on the group.
We covered the Navier-Stokes fight on September 9th, and the shape of this is the same. “More than 100” is one company’s count of its own unreleased model’s output. There are no papers, no referees and no list of which problems. A body assembled to review and communicate results, standing up after the results were announced, is a real improvement over nothing. It is not the same thing as verification, and the roster of names does not make it one.
Anthropic ships Claude Opus 5.5 at 40% lower cost and 30% faster output
Anthropic just shipped Claude Opus 5.5, and the headline is simple: same top-tier quality, noticeably lower bill.
The old Opus 5 was powerful but pricey. Opus 5.5 fixes that. It matches the performance of the bigger, more expensive Fable 5.1 model on most tasks, while costing 40% less to run on typical workloads. Output also comes back more than 30% faster.
Here is what changed under the hood:
- Pricing dropped to $4/$20 per million tokens (input/output), down from $5/$25
- Cache reads (reusing content already processed) fell 60%, from $0.50 to $0.20, which is huge for agent loops that repeat large chunks of context
- Agentic coding and computer use got major upgrades, one tester migrated a 680,000-line codebase in under a day
- Responses are cleaner: key info comes first, and it actually follows your writing instructions
One catch: four breaking changes affect existing Opus 5 integrations, so check the docs before swapping model IDs in production.
OpenAI ships GPT-6 Sol and Luna at 50% lower API prices
OpenAI just dropped two new models: GPT-6 Sol for complex, multi-step work and GPT-6 Luna for high-volume, well-defined tasks like summarizing docs or answering quick questions. Both are trained with the same techniques as the flagship Astra, just faster and cheaper.
The big story is price. These are permanent rates, not promos:
- Sol: $2 / $10 per million input/output tokens (down from $4 / $20)
- Luna: $0.10 / $0.50 per million tokens (down from $0.20 / $1.20)
- Both share a 1M token context window with 128k output tokens, no long-context surcharge
- Cached inputs now get a 90% discount, and you can swap tools or change settings without losing that cache
Sol makes half as many mistakes as its predecessor. Luna matches GPT-5.6 Sol at roughly 1/100th the cost at higher effort. You can call them via gpt-6-sol and gpt-6-luna in the API right now.
Study finds em-dash use in medical preprints nearly tripled after ChatGPT launched
Researchers analyzed 69,632 medical preprints and found something quietly telling about how AI is changing scientific writing.
The em-dash (this thing: –) used to appear in about 4% of research papers. After ChatGPT launched, that number climbed to nearly 20% by 2025. AI models love em-dashes. They use them constantly. And it shows up in the data.
The key findings:
- Em-dash use tripled across thousands of papers after AI tools became mainstream
- The rise was gradual, not a sudden spike, suggesting slow adoption of AI writing tools
- The pattern held across every test they ran, making it hard to dismiss as noise
- Boilerplate sections showed almost no change, meaning the signal is real
This does not prove any specific paper used AI. It just shows that writing patterns shifted at scale right when AI tools arrived. Think of it as a fingerprint on a population, not on any one person.
Why care? If you build tools that detect AI content, this is exactly the kind of signal worth tracking.