{"repo":"InternScience/SciEvalKit","free":true,"listed":false,"github":"https://github.com/InternScience/SciEvalKit","clone":"git clone https://github.com/InternScience/SciEvalKit.git","description":"A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language models across the full research workflow.","language":"Python","stars":85,"topics":["agent","ai","ai4science","code-generation","evaluation","evaluation-framework","gemini","gpt","llm","llm-evaluation"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"&nbsp;SciEval ToolKit A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language models across the full research workflow. &#160; &#160; &#160; &nbsp;Welcome to the official repository of SciEval ! &nbsp;Why SciEval? SciEval is an open‑source evaluation framework and leaderboard aimed at measuring the scientific intelligence of large language and vision–language models. Although modern frontier models often achieve 90 on general‑purpose benchmarks, their performance drops sharply on rigorous, domain‑specific scientific tasks—revealing a persistent general‑versus‑scientific gap that motivates the need for SciEval. Its design is shaped by following core ideas: - Beyond general‑purpose benchmarks ▸ Traditional evaluations focus on surface‑level correctness or broad‑domain reasoning, hiding models’ weaknesses in realistic scientific problem solving. SciEval makes this general‑versus‑scientific gap explicit and supplies the evaluation infrastructure needed to guide the integration of broad instruction‑tuned abilities with specialised skills in coding, symbolic reasoning and diagram understanding. - End‑to‑end workflow coverage ▸ SciEval spans the full research pipeline—such as image interpretation, symbolic reasoning, executable code generation, and hypothesis generation —instead of isolated subtasks. - Capability‑oriented & reproducible ▸ A unified toolkit for dataset construction, prompt engineering, in","default_branch":null,"files":null,"tree":[],"storefront":"/r/InternScience","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/InternScience/SciEvalKit/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}