Open · 2 seats
Red-Teaming Computer-Using Agents
Build an automated red-teaming pipeline that discovers how computer-using agents can be tricked by malicious web pages, pop-ups, and documents — then turn the findings into a public benchmark.
PhD mentorXXX
What you'll do
- Design attack scenarios across browsers, desktop apps, and documents
- Implement an attacker agent that generates and mutates adversarial environments
- Run large-scale evaluations on open and commercial CUAs
We're looking for
- Strong Python; comfortable with Playwright or Selenium
- Curiosity about how LLM agents fail
Nice to have
- Experience with security CTFs or web development