1. 22 Sep 2026

    AI: AI Loops 101. What a Gigawatt of Data Center Buys. AI-RTZ #1217 - Michael Parekh

    Post-training is the loop that turns a raw model into a product. Reinforcement learning from human feedback and from AI feedback, fine-tuning, ...

    michaelparekh.substack.com ↗
  2. 21 Sep 2026

    Meet Jev, the new AI model built to make decisions instead of generating text

    Subscribe to see fewer ads. Almedia not only helped build ChatGPT but also invented reinforcement learning from human feedback (RLHF), a post-training ...

    indianexpress.com ↗
  3. 21 Sep 2026

    The Post-training Process OpenAI Used for ChatGPT - O'Reilly

    Supervised fine-tuning (SFT) on human demonstrations; Training a reward model on human preference comparisons; Reinforcement learning with human ...

    www.oreilly.com ↗
  4. 19 Sep 2026

    Alibaba, Meituan units in trouble? China antitrust probe follows Trip.com's $776 million penalty - Mint

    ... training and benchmarking company founded by Li ... His work there included post-training analysis, data synthesis and reinforcement learning.

    www.livemint.com ↗
  5. 19 Sep 2026

    SenseTime SenseNova U1.5 Brings 8B-MoT Native Unified Vision With Open Training Code

    Post-training follows a specialize-then-unify recipe. Separate reinforcement-learning experts target visual aesthetics, bilingual text rendering, ...

    pandaily.com ↗
  6. 18 Sep 2026

    Xiaomi Livestreams MiMo-V2.6 Pro and Flash RL Post-Training Dashboard - Pandaily

    Xiaomi MiMo publicly streams MiMo-V2.6 Pro/Flash reinforcement-learning post-training—steps, tokens, rewards, and cost—at mimo.xiaomi.com/rl/, ...

    pandaily.com ↗
  7. 18 Sep 2026

    Why Studios Are Finally Embracing Generative Video|a16z - BigGo Finance

    fal engineers describe post-training the open-weight Minimax H3 video model with reinforcement learning and kernel-level optimization to reach ...

    finance.biggo.com ↗
  8. 18 Sep 2026

    Xenon Unveils AI That Can Operate Computers with 397B Parameters - BigGo Finance

    ... large language model. The model, which ranked second globally in a high-difficulty screen recognition benchmark, completed post-training using ...

    finance.biggo.com ↗
  9. 18 Sep 2026

    HiDream Unveils HiDream-O1-Video-1.0, a Native Omnimodal Video Model Built for ...

    During post-training, HiDream uses Diffusion Reinforcement Learning and a multimodal reward model aligned with human perception and aesthetic ...

    markets.financialcontent.com ↗
  10. 17 Sep 2026

    Xiaomi Livestreams MiMo 2.6 Reinforcement Learning Training: Over $1.13 Million Burned ...

    Xiaomi is livestreaming the reinforcement learning post-training process of its MiMo-V2.6 large language model on its official website in real ...

    finance.biggo.com ↗
  11. 17 Sep 2026

    Xiaomi Publishes Live Post-Training Dashboard for Mimo 2.6 RL Run, Streams Real-Time ...

    Xiaomi launched a public, live dashboard tracking the reinforcement-learning post-training run for its Mimo 2.6 model, streaming reward curves and ...

    aiweekly.co ↗
  12. 16 Sep 2026

    Salesforce taps Nvidia to launch CRM model, rejects general AI for enterprise - CHOSUNBIZ

    ... Reinforcement Learning, which are post-training stages, based on Nvidia models whose safety has already been verified." Post training is the ...

    biz.chosun.com ↗
  13. 16 Sep 2026

    Understanding DeepSeek V4.1 Flash, DeepMind's AlphaGenome Atlas and Muse

    The Sequence Learning ... There is another useful detail: DeepSeek reports that post-training retains supervised fine-tuning, reinforcement learning and ...

    thesequence.substack.com ↗
  14. 15 Sep 2026

    Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron - Unite.AI

    For post-training, Salesforce applied Supervised Fine-Tuning and reinforcement learning with Group Relative Policy Optimization (GRPO), using ...

    www.unite.ai ↗
  15. 15 Sep 2026

    Unisound Launches U2-Flash MoE as Post-Training RSI Flagship Flash Model - Pandaily

    Supporting pieces include asynchronous agent reinforcement learning with parallel workers, multi-teacher online policy distillation across math ...

    pandaily.com ↗
  16. 14 Sep 2026

    Shengshu Technology Releases Motus2 Self-Evolving World Model for Dexterous Manipulation

    In post-training, model-based reinforcement learning converts the same value signal into policy updates while freezing prediction and evaluation ...

    pandaily.com ↗
  17. 13 Sep 2026

    Long Live the Short King: Why 4-hi HBM Wins

    ... workload. We can divide AI compute into 3 major buckets: pre-training compute, post-training/reinforcement learning compute, inference compute.

    newsletter.semianalysis.com ↗
  18. 13 Sep 2026

    The US and China are racing to build 'self-improving AI'. Here's what's at stake

    ... training through post-training. Ad ... reinforcement-learning experiments, with the resulting experience feeding back into its learning process.

    amp.scmp.com ↗
  19. 12 Sep 2026

    Negative Self-Distillation Trains LLMs by Avoiding Flaws | AI Weekly

    ... reinforcement learning baselines." The abstract publishes no per ... Post-training teams should watch whether 'learn from your own bad ...

    aiweekly.co ↗
  20. 11 Sep 2026

    Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context - Unite.AI

    Post-training follows a standard supervised fine-tuning, reinforcement learning, and on-policy distillation sequence, with substantive changes ...

    www.unite.ai ↗
Sources

Where this comes from

Artificial Intelligence — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…

401 items Polled 22 Sep, 21:25 UTC 200

Agentic AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 21:25 UTC 200

AI Agents — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 21:25 UTC 200

Data Science — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 21:25 UTC 200

Large Language Model — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…

401 items Polled 22 Sep, 21:25 UTC 200

Machine Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 21:25 UTC 200

Deep Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…

401 items Polled 22 Sep, 21:25 UTC 200

LLMs Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…

45 items Polled 22 Sep, 21:25 UTC 200

AI Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…

401 items Polled 22 Sep, 21:25 UTC 200

Representation Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…

146 items Polled 22 Sep, 21:25 UTC 200

Reinforcement Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…

401 items Polled 22 Sep, 21:25 UTC 200

Generative AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

401 items Polled 22 Sep, 21:25 UTC 200

Multimodal AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…

401 items Polled 22 Sep, 21:25 UTC 200

Mechanistic Interpretability — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…

34 items Polled 22 Sep, 21:25 UTC 200

RLHF — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

36 items Polled 22 Sep, 21:25 UTC 200

Feeds are configured through the LEARN_FEEDS environment variable, so new sources can be added without a code change. Each poll sends the stored ETag and Last-Modified headers, so an unchanged feed answers 304 and costs the publisher nothing.