MedExQA is a novel benchmark in medical question-answering, to evaluate large language models’ (LLMs) understanding of medical knowledge through explanations. The work is published in ACL2024 BioNLP workshop paper. We address a major gap in current medical QA benchmarks which is the absence of comprehensive assessments of LLMs’ ability to generate nuanced medical explanations. This work also proposes a new medical model, MedPhi2, based on Phi-2 (2.7B). The model outperformed medical LLMs based on Llama2-70B in generating explanations. From our understanding, it is the first medical language model benchmark with explanation pairs.
HF: MedExQA
MedExQA is a novel benchmark in medical question-answering, to evaluate large language models’ (LLMs) understanding of medical knowledge through explanations. The work is published in ACL2024 BioNLP workshop paper. We address a major gap in current medical QA benchmarks which is the absence of comprehensive assessments of LLMs’ ability to generate nuanced medical explanations. This work also proposes a new medical model, MedPhi2, based on Phi-2 (2.7B). The model outperformed medical LLMs based on Llama2-70B in generating explanations. From our understanding, it is the first medical language model benchmark with explanation pairs.
HF: MedExQA