
This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the All Things AI Telegram group.
If you follow AI for work, research, investing, or just to understand where the technology is heading, this weekly brief is a concise way to scan the most important launches, risks, and resources in a few focused minutes.
The week of 12th July to 18th July 2026 brought a series of model, product and workflow developments. A purported internal Zhipu letter outlined AGI ambitions and plans to open‑release GLM‑5.2 with a one‑million‑token context window; Thinking Machines released the open‑weights multimodal Inkling model; and Moonshot launched Kimi K3, a 2.8T‑parameter multimodal model with 1M context and planned open weights. Anthropic introduced Claude for Teachers, announced Claude Fable 5 would be included in Max and Team Premium plans from July 20, and Google‑backed Key Studio offered support for early‑stage AI startups; tools such as Vorflux, Gemini Notebook, and PenEcho's shared AI canvas continued the shift from AI assistance towards more autonomous software and research workflows.
Safety, governance and deployment economics were equally prominent. MCPTox testing and Warrant's policy‑gated action layer drew attention to tool‑poisoning and destructive‑agent risks in production environments, while Demis Hassabis renewed calls for stronger standards for frontier models and President Xi's first appearance at the World AI Conference signalled China's state‑level commitment to open AI development. Chamath Palihapitiya's CNBC comments on the widening cost gap between Western and Chinese model providers reinforced ongoing community debate on inference economics, and reports concerning Grok Build's handling of repository data—followed by xAI statements on zero data retention—kept privacy and developer trust in focus. Cursor's expanded model allowance and Grok Build's open‑sourcing illustrated how rapidly AI products are reaching more users, even as operational safeguards and data governance remain unresolved.
The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interests.
In April 2026, a group of U.S. AI researchers traveled to China to get a firsthand look at the country’s fast-moving AI ecosystem.
During the trip, they visited AI labs and companies across Beijing, Hangzhou, and Shanghai, meeting with teams from Alibaba, Moonshot AI, Zhipu AI, Tsinghua University, Meituan, Xiaomi, Qwen, Ant Group, and 01.AI.
Qian Chen of Silicon Valley 101 in conversation with Nathan Lambert, a prominent AI researcher known for his work on RLHF and open-source AI, joined the trip. A graduate of UC Berkeley,
Nathan previously helped build Hugging Face’s RLHF research team. He later led post-training research at the Allen Institute for AI, better known as Ai2, where he worked on open models including OLMo and Tülu. He is also the author of Interconnects, one of the most widely read independent publications covering frontier AI research and policy.
The below article reproduces a translated letter allegedly written by Jie Tang, founder of Zhipu AI/GLM, arguing that AI has entered an irreversible “great wave” toward AGI.
Its main points:
- Zhipu’s strategy is built on first-principles thinking, contrarian decisions, and long-term focus, rather than short-term commercialization.
- AI’s capability ceiling is rising from perception to reasoning, with progress centered on:
1. Long-horizon task execution
2. Fully autonomous multi-agent systems
3. Self-evolving and self-training models
- Zhipu’s two-year “Touch High” initiative will focus on these areas, alongside major investment in safety, interpretability, and governance.
- The company claims it will pursue an open ecosystem, highlighting the planned open release of GLM-5.2 under the MIT License with a one-million-token context window.
- The letter frames the race toward AGI/ASI as both a technological opportunity and a major responsibility, with Zhipu aiming to push frontier capabilities upward while making them broadly accessible.
Caveat: Bing Xu says this is a translation of an internal GLM letter found on RedNote and describes it as “purportedly” written by Jie Tang; the document’s authenticity is therefore not independently established.
via Huzefa
Prompt Loops: Why the Best AI Results Come from Iteration 🔄
The most effective AI workflows do not rely on a single prompt. They follow an iterative cycle:
Prompt → Generate → Evaluate → Refine → Repeat → Final Output 🔄
Research supports this approach. Studies show that iterative prompting can improve multi-step reasoning, enhance response quality, and produce more reliable outputs than one-shot prompting. It also enables models to self-correct, refine reasoning, and better align with user intent.
As AI agents become more capable, prompt loops are evolving from a best practice into a core design pattern for building reliable AI systems. 🛠
References:
• Iteratively Prompt Pre-trained Language Models for Chain of Thought (EMNLP 2022) 📚
• Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models (NAACL Findings 2024) 📚
• Understanding the Effects of Iterative Prompting on Truthfulness (ICML 2024) 📚

via Madhav
We were building a secure boundry system for agents.
Since MCPs and tool calls are a big attack vectors even for frontier agent, We tested multiple models across MCPtox which is a tool poisining benchmark to test security across models.
Results were crazy, on avg most models allowed 30%+ critical tool uses/commands that could cause harm and warrant blocked them.
Followup by Prames
use case here would be:
agent is touching your production databases, calendars, spreadsheets etc etc
safety gated agent actions for high-impact environments
the MCPTox bench we used is published here
via Roc Zacharias,
Question. When I moved from VPS to Mac Mini for hosting my agents, it saved me money each month and worked great. Was paying like $100/month for VPS. Now pay nothing, except electricity which I think on Mac mini is almost negligible?
Does moving from external LLM to internally run have similar cost savings per month after you pay for the hardware? Right now I just use ChatGPT with Hermes for $100/month and I haven’t run out of capacity yet. I’m starting to ramp up my usage though and wondering once I get over the $100/month ChatGPT subscription, if local LLM are more economical or if they’re used more for privacy than economics?
Could a Mac mini run LLM at all btw or no, need dedicated like $5000+ hardware?
Response from Jack
Gemma4 and a few other quantised models
Nothing beefy though
But they’ll eat up most of your memory for little practical benefit
Also could get Strix halo machine for like 2-3k that can run some decent stuff
Have you looked at the costs of the Chinese models also, for less sensitive tasks they are dirt cheap, Minimax, Deepseek, Glm etc the same usage you get on the $100 plan will be like $10-20 or less on a Chinese model
Response Just|LDA
I have Qwen 3.5 running as a back up local model. I hit my usage on ChatGPT and Claude frequently. They have me on a daily drip
via Alex
The legal filings Apple v OpenAI are wild.
If only half of the things are true, it is a death sentence for OpenAI. Why would you trust OpenAI with your corporate data if they so blatantly steal from what (was) their biggest customer at the time?
https://www.documentcloud.org/documents/28453229-apple-v-openai/
I haven't personally used this product and may have fallen for the click-baity text [respect to the founder/copywriter for the tweet content ] but if someone else can get around to trying it out incase its useful in their real life work and share their feedback here that would be much appreicated by others in the group
===
Someone claiming to have a much cheaper PitchBook at $0.125/request instead of Pitchbook's $25k/yr per seat
via Alex
On a similar note. Grok caught uploading FULL repos to xAI servers...
https://glitchwire.com/news/xais-grok-build-cli-was-uploading-entire-repositories-to-google-cloud-the-compan/
This included uploads of .env files even when you didnt opt in for providing data for training and also from EU customers sending data to the US.
Every Grok Build customer must assume their application compromised ~
On another note. Has anyone here tried Nemotron Two Tower?...
https://huggingface.co/nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16
my response to the Grok brouhaha was that
This is an unfortunate detraction from the fact that Grok Build is an awesome harness and Grok-4.5 a very fast and good model particularly for coding
Demis Hassabis publishes essay on imminent AGI, its unprecedented impact, risks, and proposal for Frontier AI Standards Body
Demis Hassabis, CEO of Google DeepMind and 2024 Nobel Prize in Chemistry winner for AlphaFold, wrote an essay stating AGI is likely a few years away with an impact 10 times that of the Industrial Revolution at 10 times the speed. He warns of risks in cybersecurity, biological, nuclear threats, and self-improving systems due to commercial and geopolitical races outpacing understanding, proposing a US-led independent Frontier AI Standards Body to test powerful models before release, enforce evaluations, and promote safety measures.
Came across a pretty interesting AI innovation program backed by Google’s AI Futures Fund. Looks like they’re offering funding, cloud credits and hands-on support to early-stage founders.
Anyone building in AI and looking for funding should definitely explore this and check if they’re eligible. Feels like a solid opportunity that might be flying under the radar right now.
In case there are K-12 educators in the US in this group or members know someone in this demographic within their network
via Anthropic
We're introducing Claude for Teachers, providing verified K-12 educators in the US free access to premium Claude capabilities, a library of teaching skills, and a direct connection to evidence-based curricula, mapped to academic standards in all 50 states.
via David An
In this interview, we dive deep into the next massive evolution of the internet: Agentic Commerce. We explore how AI agents are shifting from simple chatbots into autonomous decision-makers capable of pulling data via Model Context Protocols (MCPs) and executing real-world commercial transactions.
For those in the group who like to subscribe to many newsletters including yours truly's, a howto below on how to read them in a single, distraction‑free reading queue instead of relying on your inbox
This approach may also be useful incase you have been avoiding subscribing to newsletters because now you can read them without cluttering your inbox
Whilst I personally use Wispr Flow on macOS and Android, I wanted to share Willow Voice’s offer of free, unlimited AI dictation for Mac, Windows and iOS.
Members are more than welcome to go via Willow Voice referral link which gets you one month free of their Pro plan which gives additional features over their free plan. That is, take the one month Pro plan for free , experience the features in that plan and then downgrade to foreever free if the the features of the free plan are only what you need
https://app.willowvoice.com?ref=2M629L
Tom Blomfield, co-founder of Monzo and GoCardless and former YC General Partner, recently joined Anthropic's compute team, lending operational scaling expertise to AI infrastructure challenges. The talk outlines an "AI Loop" with sensors/data, policy and tool layers, quality gates, and learning mechanisms, urging early-stage founders to build these systems now while they still can....
Prasanna S, former Rippling co-founder and CTO, launched Vorflux AI as an autonomous "autopilot" for software engineering that handles end-to-end tasks from high-level prompts through planning, coding, testing in live environments, review, and merging without constant human oversight.
The promotional video demonstrates the UI executing complex workflows like implementing features across repos and services, including mobile emulators, while the founder explains shifting from copilot models to full autonomy as AI coding capabilities surpassed human levels in 2026 benchmarks.
Vorflux secured $15M seed funding from Y Combinator, Peak XV Partners, and prominent angels; the thread argues engineering bottlenecks have moved beyond code to planning and orchestration, offering users $200 credits to test it on their backlogs
Thinking Machines has released Inkling, the new leading U.S. open weights model, debuting at 41 on the Artificial Analysis Intelligence Index
The model is 975B total parameters, has 41B active parameters, and accepts text, image, and audio input modalities. The model is accessible via Thinking Machines’ Tinker platform API (256K context window) and weights are available on HuggingFace (1M context window).
Inkling natively supports image and audio multimodal inputs, a key differentiator among open weights models.
Matt van Horn of Last30Days skill fame with a great post summarising some of the clever things being done by Grok Build. Grok Build is also currently my daily driver
SpaceXAI open-sourced Grok Build yesterday: the CLI, agent runtime, tools, and TUI. I pointed my agent at all 1.3M lines and asked for the cleverest things inside. Tl;dr of my new article:
Cursor also doubled the included usage of Cursor models on all plans.
Response from Alex
I doubt they will open source the algo. Good move though
This move might be connected to operation bluebird: https://www.jdsupra.com/legalnews/new-bird-on-the-block-operation-6468820/
For the ones that dont know, X is being sued to release the twitter trademark given it is deemed abandoned (3 years no use).
If X doesnt use the twitter brand within end of the month, the trademark could be released and twitter.new launch
via Madhav
Hey! Saw several reports on twitter and in personal experience where frontier agents and models like codex, grok build pass on destructive commands for critical data and infrastructure
We have been working on the app that becomes the boundry between you and your agent (claude code/codex)
• You define policies
• What the agents can touch
• What you want to prevent ie deleting critical data
Warrant takes care of the rest!
The setup only takes 2-3 minutes!
Follow up to Madhav's post from Prit
our latest release takes of the problem stated in the below tweet from Tibo of OpenAI
claude/codex accidentally deletes production databases, files on the filesystem, spreadhsheet data etc
we built a way to guard against that as part of our agent and have launched it at a standalone developer tool
We’re dogfooding this live in customer’s production environments with names like Squarespace, AWS, Camp Network
It was built as part of our agent’s stack - we’ve pulled it out as a standalone devtool for agent-builders focused on safe, autonomous, agentic actions!
First of its kind, please give us feedback!
Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight AI model
Moonshot AI launched Kimi K3, featuring 2.8 trillion parameters, 1 million token context length, and native multimodal capabilities for text and images. The model achieves top-tier benchmark scores, ranking just behind Claude Fable 5 Max and GPT-5.6 Sol Max while surpassing Claude Opus 4.8 on evaluations including GDPval-AA v2, AA-Briefcase, BrowseComp, DeepSWE, and Terminal Bench. It is accessible via API at $3 per million input tokens and $15 per million output tokens.
Open weights by July 27, 2026 with vLLM and SGLang promising Day-0 support
Google renames NotebookLM to Gemini Notebook and rolls out an update giving every notebook a secure cloud computer, letting it write and execute code natively
PS: I think this is Google's most underrated product. Say what you may about Gemini the model, NotebookLM is awesome particularly if you want to deep dive over a bunch of Youtube videos
Ben and Jorvik Zhang commented on NotebookLM
My wife uses it to load in her study materials and turn them into podcasts where two people discuss the information in a relatable way, she finds its a much better way of learning.
NotebookLM can serve as a kind of vector database for multimodal materials (video, audio, photos, etc.), offering much more efficient indexing.
Chamath Palihapitiya Compares AI Model Inference Costs from Anthropic, OpenAI, Meta, xAI, Google, and Chinese Models on CNBC
Chamath Palihapitiya stated on CNBC that the cost of a 'barrel of intelligence' for AI inference is $56 from Anthropic, $26 from OpenAI, $1.50 from Meta, $1 from xAI and Google, and $0.50 from Chinese models. He highlighted a 112x price gap between the highest and lowest costs.
Boris Cherny from Anthropic observes that while top engineers achieve 10x output using Claude, most teams lag in adoption, following a predictable 4-step progression that requires targeted bottlenecks and guardrails rather than just more tokens.
Advancing steps involves enabling self-verification, automated code/security reviews, multi-agent interfaces, looping, batching, and dynamic workflows to build trust in full automation across work classes
True ROI tracking focuses on engineering hours saved for tasks that would have been done manually, shifting teams from maintenance to novel building; Anthropic is at step 3 advancing to 4.
via Vincent Chow of the South China Morning Post
Key takeaways from President Xi's speech in his first ever appearance at the World AI Conference in Shanghai:
- Started the speech by referring to his signature maxim, "great changes unseen in a century are unfolding across the world"
- Said that the world has "entered an unprecedented period of active innovation on AI technology", which means "great opportunities as well as challenges for governance”
- reaffirmed commitment to open source to promote AI "openness and win-win"
- warns against "over stretching" the concept of national security as applied to AI where one country's national security is prioritised over others
- China opposes emergence of “new historical injustices” in AI (one of the most strongly worded parts of the speech)
- China in next 5 years will provide 5000 opportunities to developing countries in "AI training and seminar programmes" and "cooperation centres" - names ASEAN, League of Arab States, African Union, CELAC, SCO and BRICS
Ramchand Kumaresan of Murai Labs published a 934-page book teaching LLM construction from scratch, covering tokenizers, attention, KV cache, MoE, RLHF, quantization, and serving through 35 hands-on projects. Each chapter includes a deliberate "break the thing" exercise to deepen understanding of failure modes, directly informed by his TamilLM development work and prior research papers.
The announcement has generated solid engagement with early buyers praising its practical depth, one noting it clarified why their own LLM project extended from 6 months to 1.5 years.
The following post promotes a podcast episode of "The Bench" featuring MiniMax AI research lead Olive Jy Song, discussing timelines to reach "Fable level" (comparable to Anthropic's Claude Fable 5 frontier model from June 2026), their M3 model's native multimodality and 1M token context, and talent as the key scaling bottleneck over compute.
It covers MiniMax's research culture including early AI agents for paper tracking, open-source model progress on Vibe Code Bench, and candid insights on 996 work culture, with timestamps highlighting predictions and respected competitors.
via Just|LDA
Anthropic is integrating its advanced Claude Fable 5 model—a Mythos-class AI optimized for complex coding, long-running agent tasks, and knowledge work—into Max and Team Premium plans at 50% usage limits starting July 20.
Pro and Team Standard subscribers will keep Fable 5 access through usage credits plus a one-time $100 credit, while the company addresses unpredictable demand by expanding capacity incrementally.
The update standardizes higher-tier inclusions after staged rollouts and temporary restrictions, aiming to reduce subscriber frustration and provide clearer plan expectations
via Alex
Amazing handwriting harness I have been playing with today:
PenEcho is a shared canvas where handwriting, equations, diagrams, and spatial context become part of the conversation.
Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them
The cover image of this newsletter via generated via the Seedream 5.0 model within the Krea tool via the following prompt
Two women walking in a Chinese garden, wearing Hanfu and flowing robes, in the style of Chinese ink painting, beautiful scenery of a Chinese fairy tale, misty with white snow covering the ground, graceful figures
Yusuf Goolamabbas
Comments