1. 22 Sep 2026

    CLPS Incorporation Completes AI-Assisted Anti-Money Laundering Review Project for a ...

    ... Large Language Model. News provided by. CLPS. Sep 22, 2026, 08:30 ET. Share ... By fine-tuning small general-purpose Large Language Models (LLMs) ...

    www.prnewswire.com ↗
  2. 22 Sep 2026

    CLPS Incorporation Completes AI-Assisted Anti-Money Laundering Review Project ... - Markets data

    By fine-tuning small general-purpose Large Language Models (LLMs), CLPS ... large-language-model-302886033.html. SOURCE CLPS. Twitter · Facebook ...

    markets.ft.com ↗
  3. 22 Sep 2026

    AI: AI Loops 101. What a Gigawatt of Data Center Buys. AI-RTZ #1217 - Michael Parekh

    Post-training is the loop that turns a raw model into a product. Reinforcement learning from human feedback and from AI feedback, fine-tuning, ...

    michaelparekh.substack.com ↗
  4. 21 Sep 2026

    The Post-training Process OpenAI Used for ChatGPT - O'Reilly

    Supervised fine-tuning (SFT) on human demonstrations; Training a reward model on human preference comparisons; Reinforcement learning with human ...

    www.oreilly.com ↗
  5. 17 Sep 2026

    Salesforce Launches Koa - Destination CRM

    The Koa reasoning model and training ... To post-train the model, Salesforce applied Supervised Fine-Tuning (SFT) and reinforcement learning ...

    www.destinationcrm.com ↗
  6. 16 Sep 2026

    Understanding DeepSeek V4.1 Flash, DeepMind's AlphaGenome Atlas and Muse

    The Sequence Learning ... There is another useful detail: DeepSeek reports that post-training retains supervised fine-tuning, reinforcement learning and ...

    thesequence.substack.com ↗
  7. 15 Sep 2026

    Salesforce, NVIDIA unveil CRM domain-specific reasoning model - CIO

    ... training corpus was designed to reflect ... Salesforce post-trained the model by applying Supervised Fine-Tuning (SFT) and reinforcement learning ...

    www.cio.com ↗
  8. 15 Sep 2026

    Build an AI-powered product tagging system with Amazon SageMaker serverless model ...

    In this walkthrough, we customize Qwen3-8B with supervised fine-tuning (SFT), then optimize it with reinforcement learning with verifiable rewards ...

    aws.amazon.com ↗
  9. 15 Sep 2026

    Build an AI-powered product tagging system with Amazon SageMaker serverless model ...

    In this walkthrough, we customize Qwen3-8B with supervised fine-tuning (SFT), then optimize it with reinforcement learning with verifiable rewards ( ...

    aws.amazon.com ↗
  10. 15 Sep 2026

    Salesforce Unveils Koa, a CRM Reasoning Model on Nvidia Nemotron | AI Weekly

    The pipeline combined supervised fine-tuning with reinforcement learning and a method called "group relative policy optimization," aimed at multistep ...

    aiweekly.co ↗
  11. 15 Sep 2026

    Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron - Unite.AI

    For post-training, Salesforce applied Supervised Fine-Tuning and reinforcement learning with Group Relative Policy Optimization (GRPO), using ...

    www.unite.ai ↗
  12. 15 Sep 2026

    5 Free Microsoft GitHub Courses to Learn Data Science and Artificial Intelligence

    Today, you can learn about large language models (LLMs), retrieval-augmented generation (RAG), fine-tuning, generative AI, tool use, and AI agents ...

    www.kdnuggets.com ↗
  13. 14 Sep 2026

    Fine-tuning medical AI can improve diagnosis but also creates privacy risks

    He leads research on the accuracy and reasoning of medical language models and on multimodal AI-assisted disease diagnosis, which draws on both text ...

    medicalxpress.com ↗
  14. 14 Sep 2026

    Fine-tuning medical AI can improve diagnosis but also creates privacy risks

    Anran Li et al, Memorization in large language models in medicine prevalence characteristics and implications, Nature Communications (2026). DOI ...

    medicalxpress.com ↗
  15. 12 Sep 2026

    DeepSeek planned to retire V4-Pro for V4.1-Flash. They backed down in 45 hours - Medium

    The engineers used supervised fine-tuning, reinforcement learning and on-policy distillation, with no algorithmic changes. The training pipeline ...

    medium.com ↗
  16. 12 Sep 2026

    Berkeley Develops Humanoid Lite - I Programmer

    They also carried out experiments, including the development of a locomotion controller using reinforcement learning ... training, fine-tuning, or ...

    www.i-programmer.info ↗
  17. 11 Sep 2026

    Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context - Unite.AI

    Post-training follows a standard supervised fine-tuning, reinforcement learning, and on-policy distillation sequence, with substantive changes ...

    www.unite.ai ↗
  18. 11 Sep 2026

    SenseNova-U1.5 unifies 8B multimodal model with native 4K output | AI Weekly

    The team says it will open-source training code covering supervised fine-tuning, reinforcement learning and multi-expert on-policy distillation.

    aiweekly.co ↗
  19. 11 Sep 2026

    Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

    Training combines large-scale fine-tuning, evolutionary algorithms, and reinforcement learning. The system builds on 2 ICLR 2026 papers. TRINITY ...

    www.marktechpost.com ↗
  20. 11 Sep 2026

    Magic Matched DeepSeek V4 Pro Base Quality for $500K, Using 50 Times Less Compute

    For base models that have not yet undergone reinforcement learning or supervised fine-tuning, bpb loss is particularly important. Before RL training ...

    www.techtimes.com ↗
Sources

Where this comes from

Artificial Intelligence — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…

401 items Polled 22 Sep, 18:24 UTC 200

Agentic AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 18:24 UTC 200

AI Agents — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 18:24 UTC 200

Data Science — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 18:24 UTC 200

Large Language Model — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…

401 items Polled 22 Sep, 18:24 UTC 200

Machine Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 18:24 UTC 200

Deep Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…

401 items Polled 22 Sep, 18:24 UTC 200

LLMs Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…

45 items Polled 22 Sep, 18:24 UTC 200

AI Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…

401 items Polled 22 Sep, 18:24 UTC 200

Representation Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…

146 items Polled 22 Sep, 18:25 UTC 200

Reinforcement Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…

401 items Polled 22 Sep, 18:25 UTC 200

Generative AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

401 items Polled 22 Sep, 18:25 UTC 200

Multimodal AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…

401 items Polled 22 Sep, 18:25 UTC 200

Mechanistic Interpretability — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…

33 items Polled 22 Sep, 18:25 UTC 200

RLHF — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

32 items Polled 22 Sep, 18:25 UTC 200

Feeds are configured through the LEARN_FEEDS environment variable, so new sources can be added without a code change. Each poll sends the stored ETag and Last-Modified headers, so an unchanged feed answers 304 and costs the publisher nothing.