Red Teaming Framework for
Misalignment of Large Language Models
This project builds a security technology foundation for detecting, evaluating, mitigating, and suppressing the misalignment of large language models (behavior that deviates from human expectations and ethics).
News
Overview
Content generated by large language models may contain misalignment such as misinformation, harmful information, discriminatory content, instruction violation, license violation, and leakage of personal information. For models obtained and modified from external sources, evaluating such risks on one's own is not easy.
This project is developing a technology foundation that supports such evaluation, the Red Teaming Framework (RTF). The RTF applies a variety of attacks to a submitted model and evaluates and visualizes its responses from the perspective of misalignment. The evaluation makes use of the evaluation data foundation and adversarial evaluation data foundation that this project builds on its own.
Pages
Project Information
| Program | Key and Advanced Technology R&D through Cross Community Collaboration Program (K Program) |
|---|---|
| Principal Investigator | Jun Sakuma (Professor, School of Computing, Institute of Science Tokyo) |
| Research period | FY2024 – FY2029 (planned) |
| Grant number | JPMJKP24C3 |
| Details | JST Project Database |