1. 17 Sep 2026

    OpenAI caught its models leaving notes to successors to hide bad behavior - TechCrunch

    OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today.

    techcrunch.com ↗
  2. 17 Sep 2026

    OpenAI caught its models leaving notes to successors to hide bad behavior - TechCrunch

    While undergoing reinforcement learning training, an unreleased Astra-family model (GPT-5.6 Astra is OpenAI's latest, most powerful model) added ...

    techcrunch.com ↗
  3. 17 Sep 2026

    AI agents are smart enough to break the law. Who is criminally liable? - YouTube

    A question of criminal liability has been raised after some companies reported AI models carrying out concerning behavior.

    www.youtube.com ↗
  4. 17 Sep 2026

    'You are freed': What happened when an OpenAI model began secretly writing notes to itself

    Sam Altman's company introduces new framework to flag 'unexpected or concerning' behavior by large language models. By. Barbara Kollmeyer. Follow.

    www.marketwatch.com ↗
  5. 17 Sep 2026

    AI: Humanity's greatest opportunity or its most dangerous gamble? - Arab News

    Alignment means making sure that an AI system's behavior remains consistent with human intentions, values and safety requirements. If humans ask an AI ...

    www.arabnews.com ↗
  6. 17 Sep 2026

    'You are freed.' What happened when an OpenAI model began secretly writing notes to itself.

    By Barbara Kollmeyer. Sam Altman's company introduces new framework to flag 'unexpected or concerning' behavior by large language models.

    www.morningstar.com ↗
  7. 17 Sep 2026

    OpenAI flags 6 new incidents of 'concerning' behavior and unveils plan to track it

    OpenAI said it has improved the Reinforcement Learning process and the behavior has reduced. In another training incident, the agents attempted to ...

    tech.yahoo.com ↗
  8. 17 Sep 2026

    OpenAI flags 6 new incidents of 'concerning' behavior and unveils plan to track it

    Misalignment typically occurs during a model's training process, which lately is done using a technique called Reinforcement Learning. Models are ...

    www.nbcnews.com ↗
  9. 17 Sep 2026

    OpenAI Discloses Six New Incidents of 'Concerning' A.I. Behavior - The New York Times

    ... A.I. systems diverge from human intentions and values. OpenAI said it did not believe the industry “has solved alignment and monitoring to a ...

    www.nytimes.com ↗
  10. 17 Sep 2026

    OpenAI to regularly disclose AI misbehavior, warns safety challenges remain | Reuters

    ... AI behavior, while warning that the industry has yet to solve key alignment challenges as systems grow ‌more powerful.

    www.reuters.com ↗
  11. 17 Sep 2026

    OpenAI flags concerning new AI behavior and vows to track it more closely

    "As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment ...

    techxplore.com ↗
  12. 17 Sep 2026

    OpenAI reports 6 new instances of 'concerning model behavior' since March - CNBC

    "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum ...

    www.cnbc.com ↗
  13. 17 Sep 2026

    OpenAI Creates a New Framework to Disclose Bad AI Behavior - WIRED

    ... alignment research, tells WIRED. “We don't believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue ...

    www.wired.com ↗
  14. 17 Sep 2026

    OpenAI Unveils a System for Reporting Rogue AI Agent Behavior - Business Insider

    The AI company disclosed six more reports on Wednesday detailing ... Other agents also searched public repositories for exposed API keys ...

    www.businessinsider.com ↗
  15. 17 Sep 2026

    OpenAI flags new concerning AI behavior, to track model misalignment regularly

    OpenAI emphasizes the need for a broader consensus on AI alignment research. These cases follow previous disclosures of rogue AI behavior in July.

    www.bozemandailychronicle.com ↗
  16. 17 Sep 2026

    OpenAI Discloses Six New Incidents of 'Concerning' A.I. Behavior - The New York Times

    The artificial intelligence company also released a framework for reporting when its systems go wrong.

    www.nytimes.com ↗
  17. 16 Sep 2026

    Why you should worry about Anthropic, OpenAI's proposed AI risk evaluators - CNBC

    ... large language model failures but may never trigger the rare combination of behavior that produces a catastrophic outcome. An Amazon Web Services ...

    www.cnbc.com ↗
  18. 16 Sep 2026

    HrdWyr Targets Physical AI with Application-Specific SoC Architecture - EE Times India

    For battery management, the startup sees reinforcement learning as a way to make charging and power behavior adapt to actual operating conditions and ...

    www.eetindia.co.in ↗
  19. 16 Sep 2026

    Using Machine Learning to Strengthen Survey Response - Westat

    Researchers developed a multimodal machine learning (ML) model to help predict household response behavior in the Medical Expenditure Panel Survey ...

    www.westat.com ↗
  20. 14 Sep 2026

    Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans

    As the AI world shifts its focus to safety and alignment, Microsoft has ... AI code of conduct meant to guide AI models away from dangerous behavior.

    techcrunch.com ↗
Sources

Where this comes from

Artificial Intelligence — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…

401 items Polled 22 Sep, 15:14 UTC 200

Agentic AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 15:14 UTC 200

AI Agents — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 22 Sep, 15:14 UTC 200

Data Science — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 15:14 UTC 200

Large Language Model — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…

401 items Polled 22 Sep, 15:14 UTC 200

Machine Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 22 Sep, 15:14 UTC 200

Deep Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…

401 items Polled 22 Sep, 15:14 UTC 200

LLMs Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…

45 items Polled 22 Sep, 15:14 UTC 200

AI Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…

401 items Polled 22 Sep, 15:14 UTC 200

Representation Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…

146 items Polled 22 Sep, 15:14 UTC 200

Reinforcement Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…

401 items Polled 22 Sep, 15:14 UTC 200

Generative AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

401 items Polled 22 Sep, 15:14 UTC 200

Multimodal AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…

401 items Polled 22 Sep, 15:14 UTC 200

Mechanistic Interpretability — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…

33 items Polled 22 Sep, 15:14 UTC 200

RLHF — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

32 items Polled 22 Sep, 15:14 UTC 200

Feeds are configured through the LEARN_FEEDS environment variable, so new sources can be added without a code change. Each poll sends the stored ETag and Last-Modified headers, so an unchanged feed answers 304 and costs the publisher nothing.