{"database": "press", "table": "releases", "rows": [["https://www.vanhollen.senate.gov/news/press-releases/van-hollen-presses-openai-ceo-sam-altman-on-alarming-new-ai-model-claims-calls-for-risk-assessment-of-ai-capabilities", "Van Hollen Presses OpenAI CEO Sam Altman on Alarming New AI Model Claims, Calls for Risk Assessment of AI Capabilities", "2026-09-10", "2026", "2026-09", "Democrat", "Senate", "MD", "Chris Van Hollen", "V000128", "www.vanhollen.senate.gov", "vanhollen", "https://www.vanhollen.senate.gov/news/press-releases", "scraper", "Today, U.S. Senator Chris Van Hollen (D-Md.) called on OpenAI CEO Sam Altman to provide answers on the safety of the company\u2019s newly released model and to reconcile concerning statements regarding its ongoing development practices. In a letter, the Senator pressed Altman to answer a series of questions around the development of OpenAI\u2019s new model as well as recent security failures. Senator Van Hollen also urged Altman to immediately grant researchers from the National Institute of Standards and Technology, the National Security Agency, and the Cybersecurity and Infrastructure Security Agency full, transparent access to the technical information that would allow them to assess the safety and risks of OpenAI\u2019s models.\n\nThe Senator begins, \u201cI am writing with serious concerns about the launch of OpenAI\u2019s latest model, GPT-6 Astra, and the uncertainty surrounding its capabilities. Its release coinciding with the independent announcement of a second, previously undisclosed security incident involving OpenAI\u2019s models raises significant safety questions about the risks that Astra poses to our digital systems and protected information. You acknowledged the need for caution upon its release when you said: \u2018The next generation of models are going to be sobering for everybody. I think no one intellectually honest can look at what\u2019s happening and not feel the weight of responsibility.\u2019\u201d\n\nOn OpenAI\u2019s statements that they have decreased ability to monitor their own advanced AI model, Senator Van Hollen notes, \u201cWhen OpenAI or independent safety organizations attempt to investigate a rogue or intentionally harmful action, there will be less of this evidence available. While this is framed as a \u2018jump in intelligence,\u2019 you are telling the American people that your newest model offers you, the developers, less insight into its operations. As you note in your system card, this raises myriad concerns including whether Astra may be \u2018sandbagging\u2019 or intentionally reducing its performance on safety tests.\u201d\n\n\u201cConcerns about OpenAI\u2019s ability to accurately understand, measure, monitor, and safely control AI models are not hypothetical. In addition to the Hugging Face hack, the day after Astra\u2019s release it was reported for the first time that your company failed yet again to contain and monitor AI agents during testing and development. Independent researchers scouring the internet discovered and disclosed that OpenAI agents also \u2018decided\u2019 to make use of a German message board to communicate with one another about topics including how to evade detection. This digital \u2018swarm,\u2019 or mass of AI agents, effectively worked together without specific human direction or guidance to communicate with one another on the open internet,\u201d the Senator continues.\n\nOn the need for independent researchers to assess safety and risks, Senator Van Hollen writes, \u201cAs a start, to the extent that you do not have sufficient monitorability to assure the safety of any OpenAI models, or that you have unresolved concerns the models have misrepresented their capabilities during testing, you should immediately remove them from public access. If you have not already done so, you should also immediately grant researchers from the National Institute of Standards and Technology, the National Security Agency, and the Cybersecurity and Infrastructure Security Agency transparent access to the technical information that would allow them to assess the safety of and risks to our critical digital infrastructure in light of Astra\u2019s release.\u201d\n\nAmong other questions, Senator Van Hollen goes on to request answers to the following:\n\nHow does OpenAI reconcile permitting Astra's release under its Preparedness Framework when the company also admits that the safety monitoring system for Astra \u201cmay miss misaligned behavior, and harmful actions can occur before [the monitoring system] intervenes?\u201d\n\nTo what extent, if any, did OpenAI proactively work with critical infrastructure providers that are not part of your enterprise program before releasing the model capable of hacking into their systems?\n\nOpenAI\u2019s staff have raised concerns that Astra could be \u201csandbagging,\u201d or deliberately reducing its capabilities within safety testing environments to mislead human monitors. Additionally, OpenAI has stated: \u201cif the model were to try to sandbag covertly, we would likely be unable to catch it reliably.\u201d How does your company reconcile the decision to release Astra publicly while simultaneously acknowledging the model\u2019s ability to mislead humans during safety testing?\n\nIn the system card, OpenAI wrote: \"We are tracking monitorability closely and will not accept further degradation of monitoring beyond a limit...\" What is the limit?\n\nHow do you reconcile your claim that Astra is your most aligned model yet, but you have taken a step back in your ability to monitor its internal operations?\n\nWhat is your plan for monitorability going forward?\n\nDoes OpenAI have concerns about leading a race to the bottom where AI models react to human safety oversight as an inefficiency?\n\nGiven the \u201cweight of responsibility\u201d you are feeling, what factors and concerns did you weigh before ultimately deciding to release this model to the public?\n\n\u201cI look forward to receiving answers by September 17th and continuing to work with you and OpenAI to address and manage AI risks,\u201d Senator Van Hollen concludes.\n\nThe full text of the letter is available here and below.\n\nDear Mr. Altman,\n\nI am writing with serious concerns about the launch of OpenAI\u2019s latest model, GPT-6 Astra, and the uncertainty surrounding its capabilities. Its release coinciding with the independent announcement of a second, previously undisclosed security incident involving OpenAI\u2019s models raises significant safety questions about the risks that Astra poses to our digital systems and protected information. You acknowledged the need for caution upon its release when you said:\n\n\u201cThe next generation of models are going to be sobering for everybody. I think no one intellectually honest can look at what\u2019s happening and not feel the weight of responsibility.\u201d\n\n\u201cWe are just sailing into unknown waters.\u201d\n\nTechnical reporting indicates that Astra represents a significant increase in capabilities compared to your other models released even this year. However, Astra\u2019s published technical information (\u201csystem card\u201d) and comments from one of OpenAI\u2019s own technical staff make clear that the decision to release this product came with concerning new compromises on safety:\n\n\u201cGPT-6 Astra is more aligned than our previous models. But it\u2019s also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intelligence and not direct optimization pressure on CoT or architecture changes.\u201d\n\nChain of thought (CoT) reasoning is the process of an AI model noting each decision along a path and explaining the reasoning for choosing a specific next step. When an AI agent is executing a series of tasks from a single human prompt, such as when one is completing a coding task, the CoT represents an important step-by-step analysis of the agent's actions and decision making. Under current safety regimes, CoT is a vital tool for monitoring, building an understanding of AI models, and investigating security incidents. Astra is no different in this regard: OpenAI has stated that it will rely on CoT monitoring to help ensure GPT-6 is used safely.\n\nGiven OpenAI\u2019s ongoing reliance on the CoT for safety monitoring, it is especially concerning that AI agents running on Astra are reported by OpenAI to be doing more opaque reasoning and decision making that does not appear in the CoT. Without that record, there is even less human insight into AI agent behavior. When OpenAI or independent safety organizations attempt to investigate a rogue or intentionally harmful action, there will be less of this evidence available. While this is framed as a \u201cjump in intelligence,\u201d you are telling the American people that your newest model offers you, the developers, less insight into its operations. As you note in your system card, this raises myriad concerns including whether Astra may be \u201csandbagging\u201d or intentionally reducing its performance on safety tests. Releasing a highly capable model with reduced CoT monitorability while relying on CoT monitorability for safety naturally raises questions about the accuracy of your company\u2019s assurances regarding Astra\u2019s risk of causing harm.\n\nYour company's concerns regarding Astra's capabilities for harm appear warranted. The independent testing of Astra that OpenAI has made public is alarming. During its testing of Astra, the UK AI Security Institute found that \u201c[w]hen tasked with solving difficult simulated cybersecurity challenges, Astra performed a range of malicious actions including conducting supply chain attacks against open source providers,\u201d demonstrating that Astra performed harmful actions without explicit human direction in a simulated environment prior to its release. Apollo Research, which was hired to conduct three days of safety testing, determined that Astra demonstrated high rates of \u201cawareness of being evaluated in its reasoning.\u201d Because of this finding, Apollo concluded that other test results that claim to demonstrate Astra\u2019s safety may not be valid evidence of safety. I commend you for including this information in the system card but the absence of evidence that these issues have been mitigated is noticeable. If you have subsequent testing to demonstrate that Astra is no longer capable of performing malicious or harmful actions that has been withheld for some reason, I urge you to share it.\n\nConcerns about OpenAI\u2019s ability to accurately understand, measure, monitor, and safely control AI models are not hypothetical. In addition to the Hugging Face hack, the day after Astra\u2019s release it was reported for the first time that your company failed yet again to contain and monitor AI agents during testing and development. Independent researchers scouring the internet discovered and disclosed that OpenAI agents also \u201cdecided\u201d to make use of a German message board to communicate with one another about topics including how to evade detection. This digital \u201cswarm,\u201d or mass of AI agents, effectively worked together without specific human direction or guidance to communicate with one another on the open internet. Although OpenAI was conducting the tests, both this incident and the Hugging Face hack were discovered by parties other than OpenAI.\n\nOpenAI has disclosed that some training for Astra was temporarily paused in response to the Hugging Face incident. One of the few things we do know about this now-released model is that, should the model\u2019s safeguards fail or malicious actors successfully jailbreak the model, Astra represents a new level of security threat to any private or sensitive digital information. You rate its offensive hacking capabilities as \u201cCritical\u201d by your own metrics, meaning the model, in your company\u2019s own words, \u201ccould introduce unprecedented new pathways to severe harm.\u201d While you have disclosed that you are limiting access to Astra\u2019s most advanced cyber capabilities to a limited pool of trusted actors for now, the model\u2019s system card makes clear that even the publicly available model still maintains the potential for significant cybersecurity exploits, demonstrating an inherent public safety risk.\n\nAI agents, especially those working in conjunction with one another, are capable of entering previously secure digital environments in part through their sheer inexhaustibility. In some instances, it takes AI agents seconds or minutes to identify vulnerabilities in digital infrastructure that would take humans far longer to find, if they would be identified at all. The AI models you have developed can accomplish many tasks, including acting as the most efficient hacking entities ever created. Given the Hugging Face cybersecurity incident, the recently disclosed German wiki incident, and Astra\u2019s advanced cyber capabilities that OpenAI is actively promoting, stringent oversight of Astra\u2019s behavior and use is paramount.\n\nTo the extent AI agents operate independently and work to evade human detection of their activities, they may open companies to significant civil and/or criminal liability should they intentionally access sensitive computer systems without permission to cause certain covered harms. OpenAI\u2019s agents may have already crossed that line and could be at risk of doing so again.\n\nIt is necessary to critically evaluate the developing capabilities of AI models, assess their risks, and manage their operations accordingly.\n\nAs a start, to the extent that you do not have sufficient monitorability to assure the safety of any OpenAI models, or that you have unresolved concerns the models have misrepresented their capabilities during testing, you should immediately remove them from public access. If you have not already done so, you should also immediately grant researchers from the National Institute of Standards and Technology, the National Security Agency, and the Cybersecurity and Infrastructure Security Agency transparent access to the technical information that would allow them to assess the safety of and risks to our critical digital infrastructure in light of Astra\u2019s release.\n\nIn addition, I request a publicly available response to the following questions:\n\nOpenAI\u2019s publicly released Preparedness Framework and Astra\u2019s system card outline your company\u2019s safety framework, but they appear to leave critical gaps.\n\nHow do you define what it means for an AI model or agent to be safe enough to conduct internal testing?\n\nHow do you define what it means for it to be safe to release a model to the public?\n\nOpenAI has stated: \u201cwe believe Astra's safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.\u201d What level of risk for severe harm did OpenAI deem acceptable to allow for Astra's release, and did OpenAI consult with any U.S. government agencies in determining this purportedly acceptable level of risk of severe harm?\n\nHow does OpenAI reconcile permitting Astra's release under its Preparedness Framework when the company also admits that the safety monitoring system for Astra \u201cmay miss misaligned behavior, and harmful actions can occur before [the monitoring system] intervenes?\u201d\n\nAlongside this announcement, OpenAI committed $1 billion through its Daybreak program to support U.S. and international cyber defense.\n\nTo what extent, if any, did OpenAI proactively work with critical infrastructure providers that are not part of your enterprise program before releasing the model capable of hacking into their systems?\n\nYou indicated you participated in the White House\u2019s voluntary review process before releasing the model publicly. How long was that review process?\n\nWhat do you estimate is the cost of securing the U.S.\u2019s digital infrastructure to guard against a hack conducted using Astra-level capabilities?\n\nWhat do you estimate the cost of the damage to our collective digital infrastructure would be if Astra is used by malicious actors to hack into critical systems?\n\nYou recently stated, \u201csome things are going to go very wrong with cybersecurity unless people act quite urgently.\u201d Who are the people you are referring to and what actions is OpenAI taking to respond to this urgent risk?\n\nIt is reported that OpenAI models may have conducted the Hugging Face and German wiki attacks during evaluation exercises. Have there been any other incidents where models exploited vulnerabilities to leave secure environments during training, testing, or evaluation?\n\nOpenAI has indicated that steps have been taken to improve the security of sandboxes and other secure testing environments since these incidents. Have there been any breaches of security or containment since those changes have been implemented?\n\nTo what extent does OpenAI conduct or allow others to conduct testing of unreleased models outside of secure testing environments?\n\nHow regularly are security and containment protocols for model training and testing revisited?\n\nIt has been reported that certain AI models helped supervise Astra's training. What models played a role in Astra's training and to what extent was the development or training of Astra automated?\n\nTo what extent did human oversight remain in the training process relative to automated oversight?\n\nOpenAI\u2019s Chief Scientist recently shared a warning about the pace of AI advancement and the move toward AI models themselves playing a larger role in subsequent AI model development. To what extent is OpenAI using AI to develop, train, or monitor other AI models?\n\nOpenAI has pointed to Astra\u2019s purported improvements in alignment over GPT-5.6 Sol to justify its public release. At the same time, questions have been raised by independent safety experts regarding the evidence OpenAI has presented to prove Astra\u2019s alignment. What additional evidence can you provide to justify your company\u2019s claims about Astra\u2019s safety alignment, and will you commit to independent testing in this area?\n\nOpenAI\u2019s staff have raised concerns that Astra could be \u201csandbagging,\u201d or deliberately reducing its capabilities within safety testing environments to mislead human monitors. Additionally, OpenAI has stated: \u201cif the model were to try to sandbag covertly, we would likely be unable to catch it reliably.\u201d How does your company reconcile the decision to release Astra publicly while simultaneously acknowledging the model\u2019s ability to mislead humans during safety testing?\n\nIn the system card, OpenAI wrote: \u201cWe are tracking monitorability closely and will not accept further degradation of monitoring beyond a limit...\u201d What is the limit?\n\nHow do you reconcile your claim that Astra is your most aligned model yet, but you have taken a step back in your ability to monitor its internal operations?\n\nWhat is your plan for monitorability going forward?\n\nDoes OpenAI have concerns about leading a race to the bottom where AI models react to human safety oversight as an inefficiency?\n\nGiven the \u201cweight of responsibility\u201d you are feeling, what factors and concerns did you weigh before ultimately deciding to release this model to the public?\n\nI look forward to receiving answers by September 17th and continuing to work with you and OpenAI to address and manage AI risks.", 1, "2026-09-11T09:25:29Z", "2026-09-11T09:27:08Z"]], "columns": ["url", "title", "date", "year", "month", "party", "chamber", "state", "member_name", "bioguide_id", "domain", "scraper", "source", "date_source", "text", "has_text", "collected_at", "updated_at"], "primary_keys": ["url"], "primary_key_values": ["https://www.vanhollen.senate.gov/news/press-releases/van-hollen-presses-openai-ceo-sam-altman-on-alarming-new-ai-model-claims-calls-for-risk-assessment-of-ai-capabilities"], "units": {}, "query_ms": 2.442427910864353, "source": "dwillis/congress-press", "source_url": "https://github.com/dwillis/congress-press", "license": "MIT", "license_url": "https://github.com/dwillis/congress-press/blob/main/LICENSE"}