Anthropic is a public benefit corporation focused on creating reliable and safe AI systems. They are seeking a Red Team Engineer to ensure the safety of their deployed AI systems by uncovering vulnerabilities and simulating sophisticated threat actors.
Responsibilities:
- Conduct comprehensive adversarial testing across Anthropic's product surfaces, developing creative attack scenarios that combine multiple exploitation techniques
- Research and implement novel testing approaches for emerging capabilities, including agent systems, tool use, and new interaction paradigms
- Design and execute "full kill chain" attacks that emulate real-world threat actors attempting to achieve specific malicious objectives
- Build and maintain systematic testing methodologies that evaluate every aspect of our systems
- Develop automated testing frameworks to enable continuous assessment at scale
- Collaborate with Product, Engineering, and Policy teams to translate findings into concrete improvements
- Help establish metrics for measuring detection effectiveness of novel abuse
Requirements:
- Experience in penetration testing, red teaming, or application security
- Experience in model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors
- Strong technical skills in web application security, including hands-on expertise with security testing tools (e.g., Burp Suite, Metasploit, custom scripting frameworks)
- Experience building custom automation, including LLM-specific testing frameworks
- A track record of discovering novel attack vectors and chaining vulnerabilities in creative ways
- A public body of work such as CVEs, blog posts, or disclosed bug bounty reports
- Strong written and verbal communication skills, with the ability to explain technical concepts to varied audiences
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
- Experience with AI/ML security or adversarial machine learning
- Understanding of AI safety considerations beyond traditional security, including modern guardrails against jailbreaks
- Experience testing API security and rate-limiting systems
- Background in testing business logic vulnerabilities and authorization bypass techniques
- Background in anti-fraud, trust & safety, or abuse prevention systems
- Familiarity with distributed systems and infrastructure security
- Familiarity with abuse detection mechanisms and the ability to engineer novel bypasses
- Adaptability to understand and build engagements around emerging threats outside your direct area of expertise