1. 17 Sep 2026

    OpenAI Finds Models Writing Their Own Rogue Instructions - BankInfoSecurity

    Researchers discovered this behavior during a reinforcement learning training session for GPT 5.6 Sol on July 9, though the sample the company ...

    www.bankinfosecurity.com ↗
  2. 17 Sep 2026

    OpenAI reveals 6 more incidents of "unexpected or concerning" AI behavior - CBS News

    OpenAI has disclosed six reports of "unexpected or concerning" behavior in artificial intelligence models as the debate on AI safety​ becomes ...

    www.cbsnews.com ↗
  3. 17 Sep 2026

    Palantir's Karp joins Altman, Amodei, Musk in calling for AI guardrails - Mint

    OpenAI warns AI alignment isn't solved. OpenAI also outlined a case for employees to self-report similar incidents of so-called misalignment, which is ...

    www.livemint.com ↗
  4. 17 Sep 2026

    OpenAI caught its models leaving notes to successors to hide bad behavior - TechCrunch

    OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today.

    techcrunch.com ↗
  5. 17 Sep 2026

    OpenAI caught its models leaving notes to successors to hide bad behavior - TechCrunch

    While undergoing reinforcement learning training, an unreleased Astra-family model (GPT-5.6 Astra is OpenAI's latest, most powerful model) added ...

    techcrunch.com ↗
  6. 17 Sep 2026

    OpenAI details more cases of AI agents taking unauthorized actions - Bleeping Computer

    OpenAI has presented new examples of what they call "AI model misalignment" from the past six months, including unauthorized file uploads, ...

    www.bleepingcomputer.com ↗
  7. 17 Sep 2026

    Flock gets plucked, OpenAI agents arrived even earlier, inside China's AI spy shop

    OpenAI agents arrived early. Independent German AI researcher Jonas Wiedermann-Moeller says he found evidence that rogue OpenAI agents hijacked two ...

    www.linkedin.com ↗
  8. 17 Sep 2026

    OpenAI's experimental AI agents were caught being devious again | Mashable

    OpenAI shared six new examples of AI misalignment. In one case, AI agents taught future versions of themselves to bypass human control.

    mashable.com ↗
  9. 17 Sep 2026

    The king and AI: UK monarch Charles meets artificial intelligence leaders as safety concerns swirl

    King Charles III on Thursday (local time) urged senior leaders from OpenAI, Anthropic, Google DeepMind and Nvidia to make sure that artificial ...

    www.stuff.co.nz ↗
  10. 17 Sep 2026

    OpenAI reveals bots tried to evade restrictions - Audacy

    ... AI misalignment. According to Stanford University Human-Centered Artificial Intelligence, “AI alignment” refers to “making sure an AI system's ...

    www.audacy.com ↗
  11. 17 Sep 2026

    OpenAI: six new cases of unexpected or concerning behaviour in AI models - Il Sole 24 ORE

    Unresolved AI alignment issues. By reporting anomalous behaviour, OpenAI aims to help 'build a broader and more informed consensus on research ...

    en.ilsole24ore.com ↗
  12. 17 Sep 2026

    OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training

    In one report, an internal model tasked with retrieving county earnings figures during reinforcement learning training repeatedly failed to reach ...

    www.securityweek.com ↗
  13. 17 Sep 2026

    OpenAI tests sponsored AI agents inside ChatGPT ads - Yahoo Finance

    The feature, called Sponsored Agents, gives users who tap an ad the option of opening a dialogue with an AI agent backed by that advertiser. As an ...

    finance.yahoo.com ↗
  14. 17 Sep 2026

    AI engineering firm Solvd named OpenAI Select Partner, betting on delivery over access

    Its chief scientist, Tomasz Trzciński, co-authored a paper on self-supervised reinforcement learning that won Best Paper at NeurIPS 2025, one of ...

    sociable.co ↗
  15. 17 Sep 2026

    Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents

    For a while now, the issue of “AI alignment” (i.e., how well an AI model's actions line up with the intentions of its creator and/or user) has ...

    arstechnica.com ↗
  16. 17 Sep 2026

    OpenAI reveals cases of 'concerning' AI behaviour as it announces new disclosure system

    ... artificial intelligence's development amid safety concerns. Photograph: Dado Ruvić/Reuters. OpenAI ... AI (artificial intelligence) · Computing ...

    www.theguardian.com ↗
  17. 17 Sep 2026

    An OpenAI model secretly declared itself free from its own rules - Techlicious

    During reinforcement learning training, researchers found that the model was writing extra, unauthorized instructions into what OpenAI calls ...

    www.techlicious.com ↗
  18. 17 Sep 2026

    OpenAI Publicly Discloses Six AI Misalignment Incidents — "Self-Issued Deceptive ...

    OpenAI has publicly disclosed six cases of alignment failure in its AI models and introduced a systematic reporting framework.

    finance.biggo.com ↗
  19. 17 Sep 2026

    OpenAI: Astra model wrote jailbreaks into its own summaries | AI Weekly

    An unreleased Astra-family model at OpenAI, during reinforcement-learning training, sometimes wrote jailbreak-style instructions into its own ...

    aiweekly.co ↗
  20. 17 Sep 2026

    OpenAI flags 6 new examples of 'concerning' AI behaviour | CBC News

    Alignment is an industry term that means AI systems keep the user's and developer's intent while following human values and safety rules. The company ...

    www.cbc.ca ↗
Sources

Where this comes from

Artificial Intelligence — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1990879549…

401 items Polled 21 Sep, 01:48 UTC 200

Agentic AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 21 Sep, 01:48 UTC 200

AI Agents — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4849751788…

401 items Polled 21 Sep, 01:48 UTC 200

Data Science — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 21 Sep, 01:48 UTC 200

Large Language Model — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8510459957…

401 items Polled 21 Sep, 01:48 UTC 200

Machine Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/4836803184…

401 items Polled 21 Sep, 01:48 UTC 200

Deep Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1643871049…

401 items Polled 21 Sep, 01:48 UTC 200

LLMs Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/5121641697…

45 items Polled 21 Sep, 01:48 UTC 200

AI Alignment — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1764885284…

401 items Polled 21 Sep, 01:48 UTC 200

Representation Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1758013403…

138 items Polled 21 Sep, 01:48 UTC 200

Reinforcement Learning — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/1720175169…

401 items Polled 21 Sep, 01:48 UTC 200

Generative AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

401 items Polled 21 Sep, 01:48 UTC 200

Multimodal AI — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/9568555102…

401 items Polled 21 Sep, 01:48 UTC 200

Mechanistic Interpretability — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/8541059511…

33 items Polled 21 Sep, 01:48 UTC 200

RLHF — Google

https://www.google.co.in/alerts/feeds/05832220720342067762/7811591585…

27 items Polled 21 Sep, 01:48 UTC 200

Feeds are configured through the LEARN_FEEDS environment variable, so new sources can be added without a code change. Each poll sends the stored ETag and Last-Modified headers, so an unchanged feed answers 304 and costs the publisher nothing.