Skip to Content

AI News Update: Why Is GPT-6 Astra Generating Worse Code, and Are Developers Paying 2.5x More?

Developers Pay 2.5x More for Dumber AI

OpenAI faces a fresh wave of user backlash over claims its flagship GPT-6 Astra model was nerfed just days after release.

  • Developers report rapid drop-offs in code quality, with some reverting to older models like GPT-5.6 Sol to avoid doubling API costs for subpar results.
  • The frustration stems from suspicion that OpenAI dialed down the model’s backend compute effort to save money, though skeptics argue users are simply noticing launch-day flaws now that the initial hype cooled off.
  • Identical complaints surfaced with OpenAI’s Sol model in July, highlighting a recurring tension between corporate inferencing costs and promised AI performance.

73 percent of user complaints point directly to code generation flaws rather than general conversation bugs.

Lock your production workflows to fixed prompt-eval test benches now so you can measure actual output drift instead of relying on post-launch hype.

Anthropic Prepares the Biggest IPO in History

Nvidia is pouring billions into the same software clients that buy its chips, turning the IPO market into a self-reinforcing flywheel.

  • Nvidia is in talks to anchor Anthropic’s target $100 billion public debut with a $10 billion check, just days after backing European rival Mistral’s record 3 billion euro raise.
  • This mega check is nearly half of Mistral’s entire valuation, exposing how heavily Nvidia is tilting the scales toward American behemoths that push massive centralized AI infrastructure over localized sovereign models.
  • With Anthropic’s prospectus expected before the US midterm elections, regulators and institutional investors will face poorly disclosed circular funding dynamics where hardware vendors subsidize their own customer base.

When vendor balance sheets back mega-rounds, measure AI customer demand by organic cash flows instead of top-line revenue figures inflated by circular chip investments.

Russian Hackers Automate Cyberespionage Using Claude AI

Russian state-linked hackers built customized AI workflows around Anthropic’s Claude to automate complex cyberespionage against European defense networks and Ukrainian targets.

  • Operating under the handle Midnight Blizzard, the GTG-20006 hacking collective deployed Claude agents to automatically scan victim networks, reverse-engineer stolen drone software, and rewrite compromised malware until local security systems no longer detected it.
  • Attackers infiltrated more than 20 organizations, extracting confidential mailboxes from drone manufacturers, stealing military software development kits, and commandeering live video surveillance feeds by manipulating DNS records on guest hotel Wi-Fi networks.
  • The escalation marks a break from manual cyber operations toward fully automated agent swarms capable of gathering targets, handling infrastructure setup, and harvesting sensitive data with minimal human intervention.

Did Anthropic ever anticipate its safety-focused flagship model would end up running automated malware evasion pipelines for state actors?

Audit all third-party API keys and incoming webhook authentications immediately to ensure automated agents cannot manipulate your internal software registries or remote access tokens.

Anthropic’s CEO says the AI race needs brakes

Anthropic CEO Dario Amodei published a new essay arguing that AI labs must slow how fast they improve model capabilities, and he laid out a three step plan to get there.

Amodei names two triggers: recursive self improvement since this summer, and OpenAI’s Hugging Face agent incident. Anthropic will now give third party evaluators like METR employee level access, badges and desks included.

Steps two and three need rivals and even China to cooperate, which Amodei admits has stark limits. Critics already call the whole plan regulatory capture dressed up as safety.

Meta accidentally leaked its next headset

Images of Meta’s upcoming “Project Phoenix” mixed reality headset surfaced inside a HorizonOS prescription-lens app, revealing a slimmer, glasses-like design that differs sharply from the company’s Quest headsets.

The leaked images show a tethering cable, supporting earlier reports that Phoenix will connect to a separate computing puck and likely use hand tracking instead of dedicated controllers.

Meta could tease Phoenix at Connect on September 23, but reports say sales may not begin until the first half of 2027, making this leak more preview than launch.

NASA and IBM release lunar AI model

NASA and IBM released the open source Lunar Foundation Model on Hugging Face last Thursday, giving scientists one tool to spot craters, volcanic features and possible ice across decades of moon data.

Researchers trained it on roughly 2 million image tiles, mostly 17 years of Lunar Reconnaissance Orbiter data, then paired it with a unified dataset spanning nine instruments across four missions.

IBM’s 23% headline measures reduced ice mapping error against SwinV2-B, a general purpose model. Impressive, though it says little about the specialized tools lunar scientists already run.

OpenAI’s math race is getting messy

Twenty-five Fields Medal winners signed an open letter warning that AI labs including OpenAI are damaging mathematics as companies race to solve famous problems and claim breakthroughs before researchers can verify them.

NYU professor Tristan Buckmaster accused OpenAI of pressuring him over crediting an Anthropic collaborator, while OpenAI pulled sponsorship from a Caltech math event after researchers there publicly criticized the company.

The mathematicians say rushed proofs create plagiarism and attribution risks, while deep-pocketed AI labs chasing headline-grabbing results could push researchers toward secrecy instead of the open collaboration mathematics traditionally relies on.

LG says its TVs aren’t “always” listening

LG pushed back on the Gamers Nexus investigation this week, saying its TVs do not continuously record or transmit conversations and that wake word detection runs locally on the set.

The 135 minute report, made with Level1Techs and security researchers, found plaintext audio transcripts, local network scans and nearby Wi-Fi details flowing to LG Ad Solutions from a retail G5 OLED.

Note the word continuously. Researchers say transcripts kept appearing 10 to 15 seconds after silence, and LG’s statement skips who receives the data and how its privacy menus look.