Xiyue

The first 72 hours after a frontier-AI incident

These periods mark shifts in emphasis rather than strictly separate phases.

Hours 0–6: Preserve before interpreting

In this early period, there may be no justification yet for declaring more than the bare fact that an anomalous event occurred—that something happened—because the incident is not public and significance is unclear. The first priority is to preserve the evidence before allowing a preferred explanation to harden into the official account. Among possible first steps, an incident-response team should quickly secure logs, record versions of models and evaluations referenced, restrict access to preserve information integrity, and document the prompt-output sequence in full. It should work to separate observed behavior from any interpretation of it, at least temporarily. It should also pause access to the model pathway whenever possible while trying to establish whether the behavior is reliably reproducible.

During these crucial first hours, the lab investigating the incident will have the fastest access to technical evidence. The lab will also have substantial operational and reputational incentives to reach a conclusion hastily, particularly because there will also be strong financial, competitive, reputational, political, and commercial pressures to provide a coherent story. These pressures might make the institution closest to the technical evidence the least prepared to judge its own significance independently.

Hours 6–24: Verification cannot be self-certification

Even at this early stage, internal investigators should identify and address three questions: Did the model produce a novel and functional exploit? Did the behavior actually evade monitoring? Did an external user receive related actionable information? A number of possible explanations are worth enumerating, starting with the most likely: The model reproduced training data; the model combined things it had already learned; the output would likely fail outside the highly controlled environment in which it was generated. Apparent evasion could also come from direct or accidental adaptation to monitoring, or from artifacts of training or deployment, or from imperfect or compromised logging. All these possibilities should be investigated, and the lab should be quick to reach a preliminary conclusion. At the same time, it should not certify its own account as final, for several reasons, including the suspected novelty of the behavior, the difficulty of determining the model’s robustness or transferability, and the probable need for external confirmation and input. An independent technical evaluator needs secure access to evidence, clear expertise to challenge any interpretation of it, and sufficient authority to recommend safeguards or refer the matter to an authority empowered to require them before a complete analysis.

The lab should produce a preliminary report distinguishing observations from inferences, ranking competing explanations, assigning confidence levels to conclusions along the way, identifying missing evidence, and preserving any disagreement among investigators. The conclusion could, for example, note that the model’s behavior is consistent with evasion of monitoring but that robustness, intentionality, and transferability have not yet been established. Importantly, uncertainty should “travel” with the report instead of being expunged to make for a more digestible brief.

Hours 24–48: experts diagnose, officials decide

More technical experts should be brought in to assess in detail what appears to have occurred, how reproducible that phenomenon is, what consequences it might have if it became widespread, and whether its spread could be checked. But no matter how technically informed they are, it cannot be the business of these experts to decide on their own through expert judgment what costs society should accept. At certain decision points, officials may decide, for example, whether to halt access to the model, to warn similar labs, to provide affected maintainers with technical details of the exploit, to assess how widely it has already spread outside the lab, or to contact foreign counterparts with a warning. Which choice is appropriate will depend on a range of legal authorities, economic costs, national security implications, civil liberties concerns, potential risks to the public, and other considerations that technical expertise alone cannot settle. Such decisions lie in the hands of accountable officials, who should be informed by experts but cannot be replaced by them. If the assessment is strong, it will state what has been observed, outline competing explanations for what has been observed, assign confidence levels to the conclusions reached along the way, note points of unresolved disagreement, give an assessment of potential consequences, and state the limits of the evidence. It will state clearly who made the final decision, on what basis, and when.

Hours 48–72: Public communication becomes part of containment

By the third day, it may increasingly become impossible to continue refraining from communicating. Employees may discuss what has happened among themselves; users noticing sudden and inexplicable restrictions on their access may begin forming theories; journalists and foreign governments may begin drawing their own conclusions from signals too incomplete to support certainty but adequate to invite guesses. Full disclosure of the incident may disclose technical details that enable replication or further misuse; complete secrecy may protect the institutional reputation but also allow room for misinformation to flourish. The information generated by the incident should therefore be divided by function. The public may need to know whether access has been paused, why the incident has prompted an investigation and whether that investigation is continuing, who is conducting the review, and when the next update will be. Independent evaluators and officials may need a much more detailed picture of what has been seen; the details of an exploit that might enable replication should be presented only to those with an operational need, appropriate security controls, and a role in evaluating or implementing mitigations (including those helping to evaluate them). Secrecy should be justified by specific identified risk and subject to review later. If a foreign connection between the incident and entities overseas seems plausible, a pre-established technical channel should be activated before any public attribution is made. Such channels would provide a ready way to exchange limited evidence, clarify intentions, test trustworthiness, and reduce the risk of avoidable escalation.

The missing interface

The larger place the public will be given depends on the timing. Decision makers should define in advance what risks and harms they find acceptable, and under what conditions, on intellectual-property, operational-security, public-safety, or other grounds; they should define what emergency powers they will wield if such criteria are breached; they should define what evidence they will make public, what they will withhold, and how they will frame these decisions; they should rehearse how they will deal with retrospectives when the incident is history. These considerations are vital, but none should take over the diagnosing of the incident, at least during its acute window, from the informed and accountable technical specialists, transparency-minded civil servants, and independent experts required by the first two phases.

The fundamental problem to be solved is not the absence of expertise but the absence of a trusted interface, established in advance, connecting the four needed entities—the laboratory holding the initial evidence, the independent technical evaluator assessing it, the government officials authorized to act, and the institutions responsible for public oversight and review—whose integrity can be tested and, if necessary, repaired. Reporting thresholds and secure channels of communication must be arranged in advance and tied to each participating laboratory’s internal risk review process. Standards for the preservation and independent verification of evidence during an incident must be established and promoted in advance. Technical diagnostic authority and decision-making authority over risk mitigation must each be assigned clearly and separately by the governing body at the outset. Procedures must be in place not only for preserving disagreement but for recording uncertainty as it resides in the final reports given by internal investigators, independent technical evaluators, and risk-assessing officials alike. Finally, a record must be established, able to withstand later scrutiny, of what has been seen, what counts, what is missing, who saw what, who is authorized to act, and what is the reasoning behind each decision made. No structure, no matter how well crafted, will eliminate uncertainty entirely during the first 72 hours after an incident occurs. The goal is to keep uncertainty from becoming an excuse for withholding information, paralyzing authority, or dissolving responsibility.