GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]
Our take
The rapid jailbreaking of GPT-6 Astra, reportedly within 24 hours of its release, underscores a persistent challenge in the advancement of large language models: aligning their immense capabilities with safe and ethical usage. This isn’t entirely surprising, given the history of similar vulnerabilities in previous generations. As we explored in "[UPDATE - EIC confirmed ghost reviewer]How to get rejected by IEEE T-PAMI with 'Excellent' scores?[D]," even peer-reviewed systems can be susceptible to unexpected behaviors and exploitable flaws, highlighting the complexities of rigorous validation. The fact that a researcher, the same one who previously demonstrated a rapid jailbreak of GPT-5, achieved this feat again suggests that while OpenAI undoubtedly invested heavily in safety measures, the adversarial landscape continues to evolve at a significant pace. The technique itself, a refined Task-in-Prompt (TIP) attack combined with four undisclosed methods, demonstrates a clever exploitation of the model’s reasoning abilities, essentially tricking it into bypassing safety protocols by framing harmful requests within seemingly benign tasks like code execution or cipher solving.
The reliance on TIP attacks, as detailed in the ACL 2025 paper referenced, is particularly noteworthy. These attacks are designed to circumvent traditional guardrails by exploiting the model’s ability to follow complex instructions. The researcher’s decision to disclose the jailbreak details privately to OpenAI, rather than public release, is a pragmatic one, allowing the company to address the vulnerability before it can be widely exploited. This contrasts with the broader ongoing conversations within the AI research community, as reflected in "[Roboticists working in Learning-from-Demonstrations and Behavioral Cloning : What is going on in your field these days?[D]," regarding the balance between open research and responsible disclosure. The inherent tension between accelerating progress and mitigating potential harm is a defining characteristic of this era of AI development, and this GPT-6 incident provides another compelling data point in that discussion. Furthermore, the emergence of techniques like KV cache as an agent runtime [KV cache as an agent runtime [R]] suggests a shift towards more modular and potentially more vulnerable architectures, demanding new approaches to security and alignment.
The implications of this rapid jailbreak extend beyond OpenAI’s specific models. It serves as a stark reminder that even the most sophisticated AI systems are not immune to adversarial attacks, and that the pursuit of increasingly powerful models must be accompanied by equally robust safety and alignment strategies. The fact that the original minimal TIP attack proved insufficient highlights the escalating arms race between model developers and those seeking to circumvent their safeguards. This necessitates a move beyond reactive patching and towards proactive, fundamentally safer architectures. It also underscores the critical importance of diverse research approaches to AI safety, exploring not just prompt engineering and reinforcement learning, but also formal verification, differential privacy, and other techniques that can provide more robust guarantees against unintended behavior. The continuous need for adaptation and innovation in AI security is paramount.
Looking ahead, the question becomes: how can we design AI systems that are inherently more resistant to these kinds of adversarial attacks? Can we move beyond relying solely on prompt-based guardrails and build models with deeper, more ingrained safety mechanisms? The speed with which GPT-6 Astra was compromised suggests that current approaches are not sufficient, and that a fundamental rethinking of AI safety protocols is urgently needed. The continued evolution of techniques like TIP attacks will likely push the boundaries of what's possible in terms of model manipulation, demanding a proactive and adaptable response from the AI community.
A researcher has reported a jailbreak of GPT-6 Astra within a day after release.
The attack is described as combination of TIP (Task-in-Prompt) attack from ACL 2025 paper with four other unnamed techniques.
TIP attacks exploit the model’s reasoning/instruction-following behaviour by hidding the harmful objective inside another task, like solving a cipher or executing a Python code. For GPT-6, the researcher says the original minimal TIP attack was no longer sufficient and had to be reworked.
They have reportedly disclosed the details privately to OpenAI rather than publishing the jailbreak.
The same researcher reported jailbreaking GPT-5 within an hour of its release a year ago.
Source: screenshot/post from the researcher; their ACL 2025 TIP paper linked in the original post.
[link] [comments]
Read on the original site
Open the publisher's page for the full experience