AI feed
Aggregated from 15 sources, polled every 3 hours and deduplicated so the same story never appears twice. Links go straight to the publisher.
post-training
26 articles mention this topic.
-
22 Sep 2026
AI: AI Loops 101. What a Gigawatt of Data Center Buys. AI-RTZ #1217 - Michael Parekh
Post-training is the loop that turns a raw model into a product. Reinforcement learning from human feedback and from AI feedback, fine-tuning, ...
michaelparekh.substack.com ↗ -
21 Sep 2026
Meet Jev, the new AI model built to make decisions instead of generating text
Subscribe to see fewer ads. Almedia not only helped build ChatGPT but also invented reinforcement learning from human feedback (RLHF), a post-training ...
indianexpress.com ↗ -
21 Sep 2026
The Post-training Process OpenAI Used for ChatGPT - O'Reilly
Supervised fine-tuning (SFT) on human demonstrations; Training a reward model on human preference comparisons; Reinforcement learning with human ...
www.oreilly.com ↗ -
19 Sep 2026
Alibaba, Meituan units in trouble? China antitrust probe follows Trip.com's $776 million penalty - Mint
... training and benchmarking company founded by Li ... His work there included post-training analysis, data synthesis and reinforcement learning.
www.livemint.com ↗ -
19 Sep 2026
SenseTime SenseNova U1.5 Brings 8B-MoT Native Unified Vision With Open Training Code
Post-training follows a specialize-then-unify recipe. Separate reinforcement-learning experts target visual aesthetics, bilingual text rendering, ...
pandaily.com ↗ -
18 Sep 2026
Xiaomi Livestreams MiMo-V2.6 Pro and Flash RL Post-Training Dashboard - Pandaily
Xiaomi MiMo publicly streams MiMo-V2.6 Pro/Flash reinforcement-learning post-training—steps, tokens, rewards, and cost—at mimo.xiaomi.com/rl/, ...
pandaily.com ↗ -
18 Sep 2026
Why Studios Are Finally Embracing Generative Video|a16z - BigGo Finance
fal engineers describe post-training the open-weight Minimax H3 video model with reinforcement learning and kernel-level optimization to reach ...
finance.biggo.com ↗ -
18 Sep 2026
Xenon Unveils AI That Can Operate Computers with 397B Parameters - BigGo Finance
... large language model. The model, which ranked second globally in a high-difficulty screen recognition benchmark, completed post-training using ...
finance.biggo.com ↗ -
18 Sep 2026
HiDream Unveils HiDream-O1-Video-1.0, a Native Omnimodal Video Model Built for ...
During post-training, HiDream uses Diffusion Reinforcement Learning and a multimodal reward model aligned with human perception and aesthetic ...
markets.financialcontent.com ↗ -
17 Sep 2026
Xiaomi Livestreams MiMo 2.6 Reinforcement Learning Training: Over $1.13 Million Burned ...
Xiaomi is livestreaming the reinforcement learning post-training process of its MiMo-V2.6 large language model on its official website in real ...
finance.biggo.com ↗ -
17 Sep 2026
Xiaomi Publishes Live Post-Training Dashboard for Mimo 2.6 RL Run, Streams Real-Time ...
Xiaomi launched a public, live dashboard tracking the reinforcement-learning post-training run for its Mimo 2.6 model, streaming reward curves and ...
aiweekly.co ↗ -
16 Sep 2026
Salesforce taps Nvidia to launch CRM model, rejects general AI for enterprise - CHOSUNBIZ
... Reinforcement Learning, which are post-training stages, based on Nvidia models whose safety has already been verified." Post training is the ...
biz.chosun.com ↗ -
16 Sep 2026
Understanding DeepSeek V4.1 Flash, DeepMind's AlphaGenome Atlas and Muse
The Sequence Learning ... There is another useful detail: DeepSeek reports that post-training retains supervised fine-tuning, reinforcement learning and ...
thesequence.substack.com ↗ -
15 Sep 2026
Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron - Unite.AI
For post-training, Salesforce applied Supervised Fine-Tuning and reinforcement learning with Group Relative Policy Optimization (GRPO), using ...
www.unite.ai ↗ -
15 Sep 2026
Unisound Launches U2-Flash MoE as Post-Training RSI Flagship Flash Model - Pandaily
Supporting pieces include asynchronous agent reinforcement learning with parallel workers, multi-teacher online policy distillation across math ...
pandaily.com ↗ -
14 Sep 2026
Shengshu Technology Releases Motus2 Self-Evolving World Model for Dexterous Manipulation
In post-training, model-based reinforcement learning converts the same value signal into policy updates while freezing prediction and evaluation ...
pandaily.com ↗ -
13 Sep 2026
Long Live the Short King: Why 4-hi HBM Wins
... workload. We can divide AI compute into 3 major buckets: pre-training compute, post-training/reinforcement learning compute, inference compute.
newsletter.semianalysis.com ↗ -
13 Sep 2026
The US and China are racing to build 'self-improving AI'. Here's what's at stake
... training through post-training. Ad ... reinforcement-learning experiments, with the resulting experience feeding back into its learning process.
amp.scmp.com ↗ -
12 Sep 2026
Negative Self-Distillation Trains LLMs by Avoiding Flaws | AI Weekly
... reinforcement learning baselines." The abstract publishes no per ... Post-training teams should watch whether 'learn from your own bad ...
aiweekly.co ↗ -
11 Sep 2026
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context - Unite.AI
Post-training follows a standard supervised fine-tuning, reinforcement learning, and on-policy distillation sequence, with substantive changes ...
www.unite.ai ↗
Where this comes from
Artificial Intelligence — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…
Agentic AI — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…
AI Agents — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…
Data Science — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…
Large Language Model — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…
Machine Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…
Deep Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…
LLMs Alignment — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…
AI Alignment — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…
Representation Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…
Reinforcement Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…
Generative AI — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…
Multimodal AI — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…
Mechanistic Interpretability — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…
RLHF — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…
Feeds are configured through the LEARN_FEEDS environment variable, so new
sources can be added without a code change. Each poll sends the stored ETag and
Last-Modified headers, so an unchanged feed answers 304 and costs the
publisher nothing.