TOPIC / MODELS

OpenAI

Reporting and analysis about OpenAI models, products and infrastructure.

13 PUBLISHED STORIESUPDATED 18 MIN AGO

Latest stories

News leads this topic. Analysis and experiments appear when they add practical context.

View all model stories →

Anthropic expands long-context controls for production Claude workloads

New controls focus on when deeper reasoning should be invoked and how teams bound cost.

ANTHROPIC11 SEPT

Google DeepMind publishes a new method for evaluating multi-step reasoning reliability

The work shifts attention from single-answer accuracy toward repeatable reasoning behavior.

DEEPMIND11 SEPT

Meta updates its open model family with stronger tool-use and coding behavior

The release narrows the gap for teams that need deployable weights and controllable infrastructure.

META11 SEPT

Hugging Face launches a reproducibility suite for open model benchmarks

The toolkit records prompts, environments and evaluator settings alongside leaderboard results.

HF11 SEPT

Microsoft adds model routing policies for mixed frontier and small-model deployments

Enterprise teams can route requests by latency, cost and governance constraints.

MICROSOFT11 SEPT

Inside Models

The recurring questions shaping this topic.

01

Reasoning

Capability vs. inference cost

18 stories
02

Open weights

Control, deployment and reproducibility

12 stories
03

Multimodal

Image, audio and native tool use

9 stories
04

Evaluation

Reliability beyond leaderboards

11 stories
05

Inference economics

Latency, routing and serving cost

8 stories
06

Model safety

Behavior, controls and deployment risk

7 stories

Beyond the news cycle

Practical experiments and deeper analysis when the model story needs more than a headline.

EXPERIMENT / MODELS

Testing local models for structured tool calling under latency pressure

Same tool schema, same evaluation set, different model sizes and serving setups. We measured reliability before speed.

QWEN · LLAMA · 500 TOOL CALLS · TEST LOG 0027
92% valid calls best local configurationSee the experiment →

DEEP DIVE / INFRASTRUCTURE

The inference stack is becoming the product

Why routing, caching, observability and evaluation are moving from supporting infrastructure into the core AI experience.

14 MIN READ · UPDATED 10 SEP 2026

ANALYSIS

Benchmarks are easy. Reliable agent evaluation is not.

TOPIC CURATION

Models includes reporting, experiments and analysis about foundation models, reasoning systems, open weights, evaluation and model-serving economics. Stories are grouped by editorial relevance, not only by tags.

LAST REVIEWED · 11 SEP 2026