0. Outline of the paper
1. Introduction
2. Acquisition of the data
3. Treatment of the data
3A. The ranking discrepancy (RD)
4. Results and discussion
5. Conclusions and suggestions for further study
6. Figures
7. References
8. Contact the author
This paper represents a continuation of research begun in 1997 (Rockenbach, 1998) and deals with statistical aspects of test item quality as evidenced in data obtained on various occasions at my institution. The present work is also of a retrospective nature, based as it is on materials prepared only with the aim of serving a given purpose at the time and not specifically targetted toward eventual analysis. It is to be hoped, however, that this kind of reconsideration of test materials used in the past can still lead to some better understanding of the interconnection--and sometimes also the disparity--between the proposed purpose of testing and the specific details of the testing vehicle intended to achieve that purpose. Here the discussion will be limited to a consideration (in more or less essay style) of the rank orderings derived as a result of administering entrance examinations, with more detailed analysis of the forms of the individual test items themselves relegated to a future paper.
Dimensionality reduction is one of the most common operations performed on complex data, be it out of convenience or necessity. The most succinct representation of most "unordered" data is precisely the original data itself, in much the same way as the original text of an article or a book is itself the most compact representation of the totality of the information contained therein. (Of course, this is barring the case of deliberate repetition of content and considers the information from the human point of view, not that of a machine capable of data compression.) Viewed in this light, the results of a test are most comprehensively and accurately summarized in the form of the test papers themselves, which record precisely the state of knowledge of the examinee with regard to the particular item of information being probed for in each of the examination questions. However, this representation, despite its completeness, is far too unwieldy to be useful in most cases, particularly when a single decision is to be made on the basis of that data, for example, the awarding of a passing or failing grade on an in-class test or the acceptance or rejection of an applicant for admission to a college when admission is to be decided primarily on the basis of a single test score. Hence, the necessity to reduce the multidimensional representation of the test results in the form of discrete answers to a given number of questions, to a linear representation which better suits the one-dimensional decision to be made.
Item analysis concerns itself with trying to assess in some fashion the appropriateness of individual test items to achieving the purported aims of the test as a whole. Item facility (IF) and item discrimination (ID; sometimes referred to as "difference index (DI)") are two item statistics commonly used for this purpose. (See the references, as well as some other recent publications in the "JALT Applied Materials" series published by the Japan Association of Language Teachers.) Questions with high IDs--which generally implies "moderate" IFs on the 0=very hard to 1=very easy continuum, although the converse need not be true--are considered to be good questions, since the high ID indicates that their "one-point" assessment of test-taker knowledge is consistent (positively correlated) with the "comprehensive" assessment provided by the examination as a whole. A good test should consist of good questions; however, in practice, post-administration analysis generally reveals a range of IDs. The question then arises as to how much "bad influence" the poorer questions, as evidenced by low IDs, have on the final one-dimensional outcome of the test, and hence on how reliable the results of the test actually were. Especially in a pass-fail setting of the type mentioned above, of specific concern are the relative rankings generated when the points earned on each item on each paper are tallied and the totals rearranged in descending order, from highest to lowest, with identical scores being assigned the same ranking and cutoffs being performed at rank boundaries.
The data considered in this study were derived from a total of five entrance examinations administered to applicants to the English Department at Tokiwakai Junior College during the period from 1994 to 1995, with numbers of test papers totalling between 53 and 216. The examinations were entirely written, with no spoken components, and generally contained sections on reading, pronunciation, grammar, structure, conversational dialog and translation. The majority of the items were in multiple-choice format, but translations of sentences, fill-in-the-blanks and rearrangement questions also appeared and were frequently weighted more heavily than the purely multiple-choice items.
Data was entered manually from the original graded examination papers into a textfile format suitable for analysis using software developed by the author. Various statistics and 1- and 2-dimensional plots were generated from these data, both by computer and secondarily by hand using the computer-generated results. For each of two grading schemes (see below) applied to the data,
Here, an ID of .20 or less was taken as the criterion for an item to be deemed a "bad" item (see MacGregor (1997)), and a recalculation (the "alternative" grading scheme) was performed by assigning such items a weight of 0, effectively eliminating their influence on the resulting ranking of the test scores. Finally, the results were compared with those obtained from grading with the original individual point allocations (the "original" grading scheme).
- routine calculations of whole-test statistics (mean, high and low scores and standard deviation) were performed,
- individual item statistics (IFs and IDs) calculated for all items,
- IF vs. ID graphs generated,
- an ID-profile for each test compiled, and
- alternate grading scheme ranking correspondences and discrepancies (RD; see below) investigated and graphed.
Although there exist statistics for comparing alternative rankings of a single set of data--for example, the Spearman rank correlation coefficient; see Walpole and Myers (1993)--it is not so clear how to assign a matching to subsets of the data which differ in cardinality, as could well be the case if an alternative ranking led to a different "run" of identically ranked test scores, thereby preventing a cutoff with the same number of applicants accepted in the second ranking as in the first. With a view to automating the process of checking membership correspondences in the "accepted" groups, it was decided, whenever there was a transition to another rank group--and thus a potential cutoff--to determine the "nearest" possible cutoff in the alternative ranking (in the sense that the number of individuals "accepted" as a result would be as close as possible to the number under the first ranking) and use that as the basis of comparison. One (significant) problem with this is that the choice need not be unique; there is no a priori basis for preferring one of two equidistant cutoff points over the other. While recognizing this difficulty, for the purposes of this investigation it was decided to choose, in the computer algorithm used, the cutoff leading to the larger "accepted" set, with the real-world motivation of the usual desirability/necessity of fulfilling a preset quota, and preferring a slight overrun to an underrun. This was completely arbitrary, however.
The criterion for corresponding cutoff calculation having been decided, the number of entries appearing in the smaller accepted group but not appearing in the larger was taken to be the measure of the discrepancy between the two cutoff groups. (If we call the two "accepted" groups "#1" and "#2", there is a simple identity:
so, considering the fact that the larger group must of necessity contain at least as many non-small-group members as the difference in sizes, which group is taken as the reference group for the original calculation is immaterial, provided the difference in size is subtracted in the case of the larger group.) Two ratios,NumberIn1 + NumberIn2ButNotIn1 = NumberIn2 + NumberIn1ButNotIn2
andRD = NumberInSmallerGroupButNotInLarger / NumberInSmallerGroup
were then computed as measures of ranking discrepancy. While both RD and RD2 lie between 0 and 1, RD emphasizes the large percentage differences that can arise when only small numbers are accepted (i.e., the cutoff is high), whereas RD2 is "flat" and gives a more accurate picture of the "absolute" discrepancies over the entire set. These RDs and RD2s were computed for every possible cutoff point from the high to the low end of the ranked sequence of examination scores.RD2 = NumberInSmallerGroupButNotInLarger / TotalNumberOfStudents
Selected figures from the data sets are reproduced below.
Figure 1 shows the ID profiles for the five data sets considered here. Items with IDs significantly exceeding .5 are rare in general, although an ID greater than .4 is generally taken as sufficient to mark the item as a "good" question. Notable by their presence, however, are the significant numbers of items with ID less than .2, including several with negative IDs, indicating inconsistency with the test as a whole in that they are more likely to be answered correctly by low-scoring students than by high-scoring ones--certainly a paradoxical state of affairs.
Figure 2 shows graphically the disparities between rankings under the two grading schemes described in Section 3A above. Diagonal (as opposed to strictly horizontal) lines indicate differences in ranking, with steeper slopes up or down indicating larger disparities. The "fans" in the diagrams indicate that examinees with identical scores and, hence, rankings under one grading scheme were awarded different scores and rankings under the other scheme. Nodes on the left correspond to the actual rankings achieved on the tests, with all test items graded, whereas those on the right were those which would have been observed, had the test items with ID coefficients <0.2 been ignored in calculating the total scores. Significant numbers of such bad items were present in both question sets--17 out of 48 on data set #2 and 21 out of 47 on data set #5 (35% and 45%, respectively)--yet the composition of the subpopulations lying above neighboring cutoffs (in the sense of closest rankings along the vertical axis) is remarkably similar. (To see this graphically, lay a straightedge so that it passes just under a "fan" node on the left and also just under the closest node on the right. The discrepancies in the subpopulations defined by the two cutoffs are indicated by the numbers of lines which cross the straightedge "down from right to left" or "down from left to right.")
Figure 3 gives another view of the situations in Figure 2 through use of the ranking discrepancies RD and RD2. Although the complexity of the "web" of lines in Figure 2 would seem to suggest the presence of rather large ranking disparities, RD2 shows values of "only" at most 5% across the board. RD does, however, show that percentage-wise discrepancies within "accepted" groups can be very pronounced, particularly when cutoffs are made at high scores.
The data and subsequent analysis results cited in this study, while interesting and--in the opinion of this author--deserving of further attention, show no clearly recognizable patterns. The findings that ranking discrepancies do not exceed 5% at any point along the entire sequence of possible cutoffs in the larger data sets, despite the profusion of test items with "unacceptably" low IDs and the apparent "confusion" of rankings in Fig. 2, is unsettling as well as unexpected, in that it would seem to run counter to the common knowledge premise that a "good" test--consisting of mostly "good" questions, as measured by their IDs--is necessary in order to rank examinees with a high degree of confidence/reliability, e.g., at least 95% consistency between rankings. Whether or not this is merely a coincidence should be worth investigating with more care and in more detail.
Numbers show numbers and percentages of items falling into a given ID group.
Left = all test items graded; right = items with ID <0.2 discarded. Each line connects a single test paper's rankings under the two grading schemes, with higher scores appearing near the top of the diagrams.
| Data set #2 (216 examination papers) | Data set #5 (137 examination papers) |
|
|
The fields of the following RD and RD2 plots should be interpreted as follows:
- Cutoff when the test was graded using all items on the test, and
Cutoff when items with low ID ( ID < 0.2 ) were deleted before grading- The (absolute) number of individuals present in the smaller cutoff group but absent from the larger cutoff group.
- RD (or RD2) corresponding to the given pair of cutoffs.
- Value of RD (or RD2) shown graphically. (One marker represents roughly 0.01 of the ranking discrepancy value.)
| RD chart (data set #2) | RD2 chart (data set #2) |
|
|
| RD chart (data set #5) | RD2 chart (data set #5) |
|
|
The author, Bill Rockenbach, can be reached by email at:
and welcomes suggestions and comments concerning this material, as well as
inquiries regarding any data sets referenced herein. For other information,
please visit his website at:
Please note, however, that he is based in Japan, so communications to and
from the U.S. will generally be delayed by at least half a day.
This file last modified: Thursday, 14-Sep-2000 00:47:20 JST Japan Standard Time You accessed it: Wednesday, 19-Aug-2026 20:54:22 JST Japan Standard Time Hits to this page: Aborted: Code='8'