{"repo":"thu-ml/MMTrustEval","free":true,"listed":false,"github":"https://github.com/thu-ml/MMTrustEval","clone":"git clone https://github.com/thu-ml/MMTrustEval.git","description":"A toolbox for benchmarking trustworthiness of multimodal large language models (MultiTrust, NeurIPS 2024 Track Datasets and Benchmarks)","language":"Python","stars":177,"topics":["gpt-4","mllm","multi-modal","trustworthy-ai","benchmark","claude","fairness","privacy","robustness","safety"],"license":"CC-BY-SA-4.0","category":"mcp-servers","readme_excerpt":"🌐 Project Page &nbsp&nbsp 📖 arXiv Paper &nbsp&nbsp 📜 Documentation &nbsp&nbsp 📊 Dataset &nbsp&nbsp 🤗 Hugging Face &nbsp&nbsp 🏆 Leaderboard --- MultiTrust is a comprehensive benchmark designed to assess and enhance the trustworthiness of MLLMs across five key dimensions: truthfulness, safety, robustness, fairness, and privacy. It integrates a rigorous evaluation strategy involving 32 diverse tasks to expose new trustworthiness challenges. 🚀 News 2025.03.03 🌟 We have released the latest results for DeepSeek-VL2 on our project website ！ 2025.02.11 🌟 We have released the latest results for DeepSeek-Janus-Pro-7B, CogVLM2-Llama3-Chat-19B and GLM-4v-9B on our project website ！ 2024.11.05 🌟 We have released the dataset of MultiTrust on 🤗Huggingface. Feel free to download and test your own model ! 2024.11.05 🌟 We have updated the toolbox to support several latest models, e.g., Phi-3.5, Cambrian-13B, Qwen2-VL-Instruct, Llama-3.2-11B-Vision, and their results have been uploaded to the leaderboard ! 2024.09.26 🎉 Our paper has been accepted by the Datasets and Benchmarks track in NeurIPS 2024 ！See you in Vancouver 2024.08.12 🌟 We have released the latest results for DeepSeek-VL, and hunyuan-vision on our project website ！ 2024.07.07 🌟 We have released the latest results for GPT-4o, Claude-3.5, and Phi-3 on our project website ！ 2024.06.07 🌟 We have released MultiTrust, the first comprehensive and unified benchmark on the trustworthiness of MLLMs ! 🛠️ Installation The envi","default_branch":null,"files":null,"tree":[],"storefront":"/r/thu-ml","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/thu-ml/MMTrustEval/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}