Screening for technical flaws in multiple-choice items. A generalizability study.
Construction errors in multiple-choice items are quite prevalent and constitute threats to test validity of multiple-choice tests. Currently very little research on the usefulness of systematic item screening by local review committees before test administration seem to exist. The aim of this study was therefore to examine validity and feasibility aspects of review committee screening for item flaws. We examined the reliability of item reviewers’ independent judgments of the presence/absence of item flaws with a generalizability study design and found only moderate reliability using five reviewers. Statistical analyses of actual exam scores could be a more efficient way of identifying flaws and improving average item discrimination of tests in local contexts. The question of validity of human judgments of item flaws is important - not just for sufficiently sound quality assurance procedures of tests in local test contexts - but also for the global research on item flaws.
Forfatteren(-erne) og Dansk Universitetspædagogisk Tidsskrift/Dansk Universitetspædagogisk Netværk. The author(s) and the journal share the Copyright
Articles published in Dansk Universitetspædagogisk Tidsskrift (Danish Journal of Teaching and Learning in Higher Education) may be used (downloaded) and reused (distributed, copied, cited) for non-commercial purposes with reference to the authors and publication host.
Articles submitted to Dansk Universitetspædagogisk Tidsskrift (Danish Journal of Teaching and Learning in Higher Education may not be submitted to - or published in - other journals. Articles may be uploaded in institutional repositories if the author is required to so as part of a grant or institutional requirement.