AI feed
Aggregated from 15 sources, polled every 3 hours and deduplicated so the same story never appears twice. Links go straight to the publisher.
supervised fine-tuning
12 articles mention this topic.
-
21 Sep 2026
The Post-training Process OpenAI Used for ChatGPT - O'Reilly
Supervised fine-tuning (SFT) on human demonstrations; Training a reward model on human preference comparisons; Reinforcement learning with human ...
www.oreilly.com ↗ -
17 Sep 2026
Salesforce Launches Koa - Destination CRM
The Koa reasoning model and training ... To post-train the model, Salesforce applied Supervised Fine-Tuning (SFT) and reinforcement learning ...
www.destinationcrm.com ↗ -
16 Sep 2026
Understanding DeepSeek V4.1 Flash, DeepMind's AlphaGenome Atlas and Muse
The Sequence Learning ... There is another useful detail: DeepSeek reports that post-training retains supervised fine-tuning, reinforcement learning and ...
thesequence.substack.com ↗ -
15 Sep 2026
Salesforce, NVIDIA unveil CRM domain-specific reasoning model - CIO
... training corpus was designed to reflect ... Salesforce post-trained the model by applying Supervised Fine-Tuning (SFT) and reinforcement learning ...
www.cio.com ↗ -
15 Sep 2026
Build an AI-powered product tagging system with Amazon SageMaker serverless model ...
In this walkthrough, we customize Qwen3-8B with supervised fine-tuning (SFT), then optimize it with reinforcement learning with verifiable rewards ...
aws.amazon.com ↗ -
15 Sep 2026
Build an AI-powered product tagging system with Amazon SageMaker serverless model ...
In this walkthrough, we customize Qwen3-8B with supervised fine-tuning (SFT), then optimize it with reinforcement learning with verifiable rewards ( ...
aws.amazon.com ↗ -
15 Sep 2026
Salesforce Unveils Koa, a CRM Reasoning Model on Nvidia Nemotron | AI Weekly
The pipeline combined supervised fine-tuning with reinforcement learning and a method called "group relative policy optimization," aimed at multistep ...
aiweekly.co ↗ -
15 Sep 2026
Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron - Unite.AI
For post-training, Salesforce applied Supervised Fine-Tuning and reinforcement learning with Group Relative Policy Optimization (GRPO), using ...
www.unite.ai ↗ -
12 Sep 2026
DeepSeek planned to retire V4-Pro for V4.1-Flash. They backed down in 45 hours - Medium
The engineers used supervised fine-tuning, reinforcement learning and on-policy distillation, with no algorithmic changes. The training pipeline ...
medium.com ↗ -
11 Sep 2026
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context - Unite.AI
Post-training follows a standard supervised fine-tuning, reinforcement learning, and on-policy distillation sequence, with substantive changes ...
www.unite.ai ↗ -
11 Sep 2026
SenseNova-U1.5 unifies 8B multimodal model with native 4K output | AI Weekly
The team says it will open-source training code covering supervised fine-tuning, reinforcement learning and multi-expert on-policy distillation.
aiweekly.co ↗ -
11 Sep 2026
Magic Matched DeepSeek V4 Pro Base Quality for $500K, Using 50 Times Less Compute
For base models that have not yet undergone reinforcement learning or supervised fine-tuning, bpb loss is particularly important. Before RL training ...
www.techtimes.com ↗
Where this comes from
Artificial Intelligence — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…
Agentic AI — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…
AI Agents — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…
Data Science — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…
Large Language Model — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…
Machine Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…
Deep Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…
LLMs Alignment — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…
AI Alignment — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…
Representation Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…
Reinforcement Learning — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…
Generative AI — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…
Multimodal AI — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…
Mechanistic Interpretability — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…
RLHF — Google
https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…
Feeds are configured through the LEARN_FEEDS environment variable, so new
sources can be added without a code change. Each poll sends the stored ETag and
Last-Modified headers, so an unchanged feed answers 304 and costs the
publisher nothing.