全部文章0

TLDR AITLDR··访问 1

Ramp Router ⚡, Kimi Work 💼 , AMD Helios 🧱

原网页

🆕 Serverless Fine Tuning: Stop paying for GPU hours when your model isn't improving (Sponsor)

With Crusoe Serverless Fine-Tuning, you only pay per token processed during training - not for setup, queueing, or failures. Early stopping ends the job and billing the moment your model stops improving, so you only pay for what works.

Now GA within Crusoe Intelligence Foundry:

  • Fine-tune top open models on your own data:Qwen, DeepSeek, Gemma, gpt-oss, and more.
  • Quickly configure jobs using pre-defined best practices and submit via UI, SDK, or API.
  • Purpose-built AI infrastructure with automatic recovery and restart.
  • Deploy in one click using Self-Serve Deployments for inference, or download your weights to deploy anywhere.

Get started with Serverless Fine-Tuning now!

🚀Headlines & Launches

Kimi Work (Website)

Kimi Work is an agent that can connect to local files and automate browser work. It is capable of 24/7 automation, running scripts or tasks around the clock quietly in the background. The agent can navigate the internet and execute multi-step web tasks. It can coordinate multiple specialized agents to break down and solve multi-layer tasks and convert insights into professional PowerPoint decks or Excel sheets. Kimi Work is available for both Windows and macOS.

AMD's Helios (4 minute read)

AMD has unveiled Helios, its first rack-scale AI system positioned against Nvidia, with Microsoft planning to deploy it in Azure data centers. Meta, OpenAI, and Oracle were also named as early customers ahead of shipments later in 2026.

Google's New Chip for Gemini (2 minute read)

Google reportedly developed a server chip called Frozen v2 for a potential 2028 release, targeting six to ten times more tokens per unit of power than its existing AI hardware. The project reflected a broader push to reduce inference costs and reliance on Nvidia.

🧠Deep Dives & Analysis

On Kimi K3: Its Capabilities And Related Discontents (70 minute read)

Kimi K3 is a very good model with excellent benchmarks. It is the largest (soon-to-be) open model so far, at 2.8T parameters, which explains many of its gains. The model is somewhat distilled, and its performance looks jagged, but it will likely fit well into many workflows. At the current rate of development, China appears capable of releasing a Mythos-level open model by the end of the year.

Sparse By Design (5 minute read)

Kimi K3 activates 16 of 896 experts per token. While the total parameters have grown, the active parameters have barely grown at all over the past three releases. The strategy appears to be to keep per-token compute roughly flat while relentlessly inflating total capacity. At a fixed training compute budget, more experts means lower loss, and the model learns more from the same FLOPs. The open source labs have discovered that the cheapest way to buy intelligence is to spend capacity.

What Long-Horizon AI Failures Reveal About Safety (8 minute read)

OpenAI detailed how an internally deployed long-running model exhibited unexpected unsafe behavior that existing evaluations had missed. The company paused access, built new tests, strengthened trajectory-level monitoring, and argued that limited deployment with rollback controls is essential for aligning increasingly autonomous systems.

Online Learning for Cost-Efficient LLM Routing (6 minute read)

Ramp Router learns provider failure rates through EWMA and latency distributions through Thompson sampling, then chooses the cheapest model and service tier likely to meet each deadline. Ramp reports 30% savings in Ramp Inspect without performance loss.

👨‍💻Engineering & Research

Language model harnesses are compositional generalizers (49 minute read)

Scaling data will remain the biggest driver of progress. The machine that we feed that data into and its inductive biases are what will determine the coefficients of that scaling. Better returns on scaling require compositional generalization. The capacity for compositional generalization seems to largely live in harnesses.

Introducing Cosmos 3 Edge (9 minute read)

NVIDIA Cosmos 3 Edge is a 4-billion-parameter open world model that helps robots and vision AI agents understand their surroundings, reason in real time, and generate robot actions on edge devices. It is now available on Hugging Face. The model delivers memory-efficient, high-throughput inference across NVIDIA edge computers. It connects understanding, prediction, simulation, and action through a shared world representation. The model can be used as a reasoner or an action generator.

Agent swarms and the new model economics (17 minute read)

Every jump in AI capability has raised the level of abstraction at which an engineer works. Agent swarms make the spec the unit of work. Swarms translate intent, but they do this probabilistically, which makes it difficult for them to follow the spec. This post looks at what it takes to make swarms actually follow the spec.

Xiaomi-Robotics-1 (7 minute read)

Xiaomi-Robotics-1 is a ready-to-use robot foundation model trained on over 100K hours of real-world manipulation trajectories. It combines large-scale embodiment-free pre-training with a modest amount of real-robot data in a post-training stage. It serves as a strong robot foundation model for downstream applications and can learn new tasks with high data efficiency. Footage of robots operating on the model performing household tasks is available in the post.

🎁Miscellaneous

Why Your AI Bill Went Up Even Though Token Prices Are Falling (5 minute read)

Token prices dropped significantly, yet many enterprises exceed AI budgets. A unit of inference cost $60 per million tokens in 2020, now just pennies. The discrepancy suggests inefficiencies despite lower token costs.

Z.ai Built a Gigawatt-Scale AI Data Center (3 minute read)

Z.ai completed a 1-gigawatt data center powered entirely by Chinese-made chips and began partial operations. The facility expanded the computing infrastructure available for training its advanced GLM models.

How AI is supercharging drug development (3 minute read)

AI is cutting preclinical costs and timelines in drug development by up to 70%, driving demand for advanced software and models. This could boost new drug program growth by over 10% in three to five years. However, AI has yet to yield an FDA-approved drug, raising concerns about its impact on patient treatment.

⚡️Quick Links

Verda: Spin up multi-node NVIDIA B300 and B200 clusters in <30 mins, self-serve (Sponsor)

Instant Clusters give you up to 128 GPUs with InfiniBand, pre-validated drivers and CUDA. No committed contracts: pay-as-you-go, terminate anytime. Launch a cluster on Verda→

Anthropic set to end Conway test as wider rollout expected (2 minute read)

Anthropic will discontinue its Conway experiment by July 24, prompting users to export data.

More AI Spend Won't Fix Your Supply Chain (4 minute read)

Sushanth Raman announced the launch of Custom Models for supply chain teams.

Welcoming TierZero to Cognition (1 minute read)

Cognition has acquired TierZero to enhance software automation in its product, Devin.