hirly

Apply with hirly

Jailbreaking Lead, Red Team

FAR.AI · Remote (International)

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

About Us FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response. We’re structured to support that work from early research through real-world adoption: Independent by design. We can pursue what's most impactful based on our theory of change and share what we find publicly. A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments. Serious infrastructure for ambitious research . A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra. Setting the standard . Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety. Since our founding in July 2022, we've grown to 50+ staff , published 40+ academic papers , and convened leading AI safety events . Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026 , and ICLR, and features in the Financial Times , Nature News , Wired Magazine and MIT Technology Review . We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the

EU AI

Office and publish the AI Security Leaderboard based on our red-teaming expertise. We help steer and grow the AI safety field through developing research roadmaps with renowned researchers such as Yoshua Bengio; running FAR.Labs , an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants to technical researchers. About Red-Teaming at FAR.AI FAR.AI’s red team is building toward a simple outcome: materially raising the bar for safety and security of the most widely deployed and capable AI systems in the world. We intend to be the tip of the spear in AI safety: the team that consistently finds the failures others miss, resulting in real mitigations, and setting the standard that labs and governments converge on. We also leverage our in-depth understanding of weaknesses in frontier models to advise frontier developers on mitigations, to guide our own research and grant-making for improving model security, and to inform the public of key AI risks. We are already one of the leading independent red-teaming organizations. Our work has helped most Western frontier model developers improve safeguards through pre- and post-deployment testing (e.g., we have directly influenced safeguards at major frontier developers like OpenAI and Anthropic), and we are increasingly embedded in high-leverage government efforts (e.g., leading a consortium building CBRN evaluations for the European Commission/EU AI Office, and collaborating with the

UK AI

Security Institute). "FAR.AI's pre-deployment testing of GPT-5 series models identified failure modes and mitigations, improving the security of our model releases." – Senior Technical Program Manager, OpenAI “FAR.AI have been a trusted and thoughtful collaborator for us, and they have progressed the state of frontier red-teaming through research like STACK. We expect this to be a high impact role and are excited to explore collaborations with the successful candidate.” – Xander Davies, Technical Lead, Red Team at

UK Aisi

You will be the senior technical owner of our jailbreaking practice reporting to Kellin Pelrine with a dotted line to Edward Yee . In 2026, we are scaling from a strong team with standout wins into a new level of impact for any AI red team globally:

  • Red-teaming all major frontier model releases (closed and open-weight) within days/weeks of release;
  • Expanding strategic engagements with governments and conducting pre-deployment testing with most frontier labs;
  • Deepening our testing of key risk areas like CBRN, cyber, and agents, and exploring new ones like AI control and alignment;
  • Building tools, agents, and insights that raise the global standard for red-teaming. About the Role Jailbreaking is the core technical engine of the red team. As Jailbreaking Lead, you own that engine. You are the person who personally breaks the hardest targets, sets the bar the rest of the team pushes toward, and makes sure we keep discovering the highest severity, universal vulnerabilities – the most important vulnerabilities to fix – in the most heavily defended frontier models on the planet, faster than anyone else. We expect you to spend at least 50-70% of your time hands-on across 2026: breaking models, chaining novel attack classes through defense-in-depth stacks, helping to invent new techniques when existing ones fail, and setting the standard for what constitutes a significant vulnerability and a credible mitigation. The remaining time will go to managing/mentoring ICs, helping to shape the jailbreaking research agenda with Kellin, and making sure our findings land with frontier labs, governments, and the broader field. The rest of the red team will empower your work, whether through direct collaboration and support, novel research and red-teaming infrastructure, or toolkits and agent build-outs. This is a senior IC role by default, intended to attract a world-class jailbreaker whose personal mission is to find critical jailbreaks in the most heavily defended domains of the leading frontier AI models, and who has a track record of repeatedly doing so. We are open to a management track for candidates who want to hire and lead a jailbreaking team over time. We will not water down the IC bar to support the management track: both versions of this role require you to be, or be on a clear trajectory to being, one of the best jailbreakers in the world. In practice, this role spans:
  • Lead jailbreaking on the highest-stakes engagements:
  • Personally develop universal and near-universal jailbreaks against frontier closed- and open-weight models, in CBRNE, cyber, agentic security, extreme persuasion, and emerging risk domains;
  • Systematically dismantle defense-in-depth stacks (input filters, model-level refusal and safe completion, reasoning monitors, output filters, account-level moderation), chaining novel and established techniques;
  • Escalate initial vulnerabilities to expose their most severe form, maximizing universality, success rate, and capability of elicited output;
  • Own the technical bar for vulnerability severity and generality on every major engagement.
  • Push the frontier of jailbreaking techniques:
  • Invent new attack classes when existing techniques fail (e.g., we have recently shipped novel attacks against Constitutional Classifiers and fine-tuning APIs);
  • Monitor and rapidly incorporate state-of-the-art methods from the literature, and build our own proprietary portfolio;
  • Shape the jailbreaking research agenda in partnership with Kellin, ensuring our toolkit stays ahead as defenses evolve;
  • Stress-test novel affordances (innovations in agents, tool use, long context, multimodal, reasoning, etc.) as frontier systems evolve.
  • Raise the technical bar across the team:
  • Set the standard for rigor, creativity, and precision in jailbreaking across the red team;
  • Mentor ICs on attack craft, running pairing sessions, post-engagement retros, and internal write-ups that turn your craft into team capability;
  • Review major red-teaming deliverables for technical quality, severity judgment, and clarity;
  • If on the management track: hire, manage, and grow a jailbreaking team without sacrificing your personal technical edge.
  • Translate jailbreaks into r