1. 22 Sep 2026

    Exclusive Interview with Jev Inventor: Keep Workflows Under Program Control – Embed AI ... - 36氪

    RLHF (Reinforcement Learning from Human Feedback) optimizes for "what humans like", RLVR (Reinforcement Learning from Verifiable Reward) focuses ...

    eu.36kr.com ↗
  2. 22 Sep 2026

    Unlocking Jev Black Box: Exclusive Dialogue with World's First Batch of Reproducers - 36氪

    This will be the next big thing after RLHF." From millisecond-response browser plugins, to high-frequency on-chain trading decision-making, to the ...

    eu.36kr.com ↗
  3. 22 Sep 2026

    TypeSafe's Diogo Almeida: The $1 Trillion Token Surge That Proves 'Code Is the New Consumer'

    Almeida contends that chat-first training methods like RLHF have fractured model intelligence, producing sycophancy and hallucinations that make AI ...

    finance.biggo.com ↗
  4. 21 Sep 2026

    Meet Jev, the new AI model built to make decisions instead of generating text

    Subscribe to see fewer ads. Almedia not only helped build ChatGPT but also invented reinforcement learning from human feedback (RLHF), a post-training ...

    indianexpress.com ↗
  5. 21 Sep 2026

    3 Major AI Giants Suffer Successive Security Setbacks: Why Large Language Model ...

    Today's safety and alignment are mainly built around four pillars. The first pillar is RLHF (Reinforcement Learning from Human Feedback). It is the ...

    eu.36kr.com ↗
  6. 21 Sep 2026

    Jev suddenly explodes and goes viral: doesn't chat, doesn't write code, only makes judgments

    More interestingly, the founder of the Jev model, Diogo Almeida, actually participated in RLHF—and now he's starting to reflect on it: this ...

    www.binance.com ↗
  7. 20 Sep 2026

    Jev sparked a weekend frenzy: A key contributor to GPT has begun reevaluating RLHF ... - 富途资讯

    Even more intriguing is that Diogo Almeida, the founder of the Jev model, was once involved in RLHF—and now he's beginning to question it: this ...

    news.futunn.com ↗
  8. 20 Sep 2026

    Viral Global AI Star Jev: The Hottest New Trending AI That Never Speaks - 36氪

    ... (RLHF) algorithm. This to some extent defined the direction of large language models over the past few years. But while this approach has driven ...

    eu.36kr.com ↗
  9. 19 Sep 2026

    Jev Makes Fast and Cheap Decisions - by Patrick McGuinness - AI Changes Everything

    Founded by former OpenAI researcher Diogo Almeida, who was a co-inventor of RLHF and ChatGPT, and his colleagues Erik Gafni and Sasha Sheng ...

    patmcguinness.substack.com ↗
  10. 18 Sep 2026

    Jev Makes Fast and Cheap Decisions - by Patrick McGuinness - AI Changes Everything

    LLM training originally relied on optimizing AI models to please human judges, with RLHF, Reinforcement Learning with Human Feedback. We have ...

    patmcguinness.substack.com ↗
  11. 18 Sep 2026

    A "silent" AI has taken social media by storm. Is Jev truly a new paradigm?

    ... RLHF training process later. OpenAI also listed him in the "Foundational RLHF and InstructGPT work" section of the GPT-4 contributor list. After ...

    eu.36kr.com ↗
  12. 18 Sep 2026

    New 'Reinforcement Learning For Calibrated Decisions' Makes AI Headlines But Look Past The Hype

    ... RLHF (reinforcement learning with human feedback). AI makers have leaned heavily into RLHF, which is partially what made ChatGPT into a great ...

    www.forbes.com ↗
  13. 18 Sep 2026

    A new kind of AI model from a ChatGPT inventor is thrilling developers | TechCrunch

    Almeida was an OpenAI researcher who helped build the chatbot and then invent reinforcement learning from human feedback (RLHF), the ...

    techcrunch.com ↗
  14. 18 Sep 2026

    TypeSafe AI's Jev Is Not an LLM — And That May Be the Point - Forkast.News

    An OpenAI Veteran's Bet Against LLMs. At the helm of TypeSafe AI is CEO Diogo Almeida, an OpenAI veteran and co-inventor of RLHF, whose work was ...

    forkast.news ↗
  15. 11 Sep 2026

    OpenAI Appoints Doomsday Theorist to Board: New Member Warns AI Could Kill Most of Humanity

    joins OpenAI's top governance layer. Christiano is one of the founders of RLHF (Reinforcement Learning from Human Feedback), the core technology of ...

    eu.36kr.com ↗
  16. 08 Sep 2026

    Opaque recurrence, and other AI terms that you should probably know | TechCrunch

    Sit in on any product meeting, pitch, or panel these days, and you'll hear people toss around LLMs, RAG, RLHF — and, as of last week, terms like ...

    techcrunch.com ↗
  17. 06 Sep 2026

    Training the Human Neural Network - RLHF to RLDF. | Ibrahim Mukherjee - The Blogs

    Repeat the process often enough and behaviour changes. In contemporary AI, RLHF normally means Reinforcement Learning from Human Feedback: humans ...

    blogs.timesofisrael.com ↗
  18. 01 Sep 2026

    How AI Data Annotation Is Powering the Next Generation of AI Models - Analytics Insight

    Human feedback helps train AI models through RLHF, where people compare and rank responses to make AI more helpful, accurate, and safer.

    www.analyticsinsight.net ↗
  19. 27 Aug 2026

    The Finance Lab Introduces TFL Bloodhound, a Financial Reasoning Model Trained on ...

    TFL_Bloodhound replaces the human evaluator in RLHF with realized market outcomes, and separates quantitative estimation from language reasoning ...

    www.digitaljournal.com ↗
  20. 26 Aug 2026

    AI Feedback Loops Explained: How Artificial Intelligence Learns, Improves

    RLHF is a technique where humans rate AI outputs and provide guidance that helps models produce higher-quality and safer responses. 4. Can AI systems ...

    www.analyticsinsight.net ↗
Sources

Where this comes from

Artificial Intelligence — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…

401 items Polled 22 Sep, 09:12 UTC 200

Agentic AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 09:12 UTC 200

AI Agents — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 09:12 UTC 200

Data Science — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 09:12 UTC 200

Large Language Model — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…

401 items Polled 22 Sep, 09:12 UTC 200

Machine Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 09:12 UTC 200

Deep Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…

401 items Polled 22 Sep, 09:12 UTC 200

LLMs Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…

45 items Polled 22 Sep, 09:12 UTC 200

AI Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…

401 items Polled 22 Sep, 09:12 UTC 200

Representation Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…

146 items Polled 22 Sep, 09:12 UTC 200

Reinforcement Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…

401 items Polled 22 Sep, 09:12 UTC 200

Generative AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

401 items Polled 22 Sep, 09:12 UTC 200

Multimodal AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…

401 items Polled 22 Sep, 09:12 UTC 200

Mechanistic Interpretability — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…

33 items Polled 22 Sep, 09:12 UTC 200

RLHF — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

31 items Polled 22 Sep, 09:12 UTC 200

Feeds are configured through the LEARN_FEEDS environment variable, so new sources can be added without a code change. Each poll sends the stored ETag and Last-Modified headers, so an unchanged feed answers 304 and costs the publisher nothing.