0. Outline of the paper
1. Introduction
2. Acquisition of the data
3. Treatment of the data
4. Results and discussion
5. Conclusions and suggestions for further study
6. Figures
7. References
8. Contact the author
In this short paper I present some preliminary findings regarding appropriateness of certain types of questions in assessing comprehension of course material by students in some of my conversational English classes. This initial attempt at test item analysis is based on a single set of data collected in September 1997 and, consequently, the results presented herein must be considered to be, at best, of a tentative nature. However, it seems reasonable to expect that similar, more detailed studies conducted on other sets of test data, particularly if the tests have been designed with a view toward eventual data gathering of this sort, might yield results useful to language teachers in improving the quality and accuracy of at least some aspects of their assessments of learner achievement.
How to assess learner achievement is an important problem faced by most educators. Accurate appraisal of the extent to which students have mastered course content is necessary not only for such tasks as the assignment of grades but also, and even more importantly, as one of the tools available for providing feedback regarding what worked and what didn't work as planned during the term, which in turn assists in the continuing quest for ever more productive and effective methods of teaching. Traditional testing relies on "point sampling" in a Monte Carlo fashion, in the sense that, by checking whether or not the learner can answer correctly a limited number of questions we hope to be able to arrive at some conception of the overall profile of that learner's mastery of a given body of knowledge. It is obviously to our advantage to know whether a particular local data point in such a scenario actually has relevance when it comes to making inferences regarding the global distribution of achievement, since a more accurate assessment would result if ineffective or marginally effective items could be replaced with more effective ones from the outset.
Simple measures of item difficulty such as the percentage of correct answers obtained are of some use; however, it is difficult to grasp the bearing which particular items have on the test as a whole when the total number is large and there is nothing more "concrete" than a decimal value to assist us in their interpretation. An excellent, thought-provoking paper by MacGregor [2] uses several other statistics to assess the quality of a well-known English proficiency test (in multiple-choice format) here in Japan, the details of which (statistics) are said to be given in another work by Brown [1]. In the paragraphs which follow, I make heavy use of these statistics and rely on the brief descriptions given in MacGregor [2] since the latter work is, at the time of this writing, out of stock at the publisher and not otherwise available to me. In particular, see the section of MacGregor [2] subtitled "Item Statistics" (pp. 29-30).
A test consisting of 49 spoken items was administered simultaneously to 90 second-year Japanese junior college students enrolled in four sections of an English conversation course. Each item counted 1 point but was graded so as to permit the possibility of 1/2 point partial credit for "sufficiently close approximations" to the correct answer(s). The items themselves were pre-recorded on a cassette to ensure uniformity of time allotted for listening to the item (repeated a uniform number of times in each section of the test) and writing the answer. Part I of the test consisted of 35 questions which were asked in English and were to be answered in English; the questions themselves were predominantly of the WH- type (29), but also included 2 yes/no and 4 choice questions, and were drawn from various materials used during the preceding semester. Part II consisted of 14 dictation-based items; each of 7 English sentences was repeated a fixed number of times and was to be dictated verbatim (7 items), after which the sentence was to be translated into Japanese (7 more items, interleaved with the preceding). The tests were graded and the individual item scores entered manually into a specially-formatted computer textfile.
The textfile containing the data was used as the input for a specially-written QuickBASIC program which processed the data and output the generated numerical quantities and graphical representations, some of which are reproduced within the body of this paper. Specifically, in addition to the basic whole-test statistics of high, low and mean score, and standard deviation, the following item statistics were computed and graphed:
1. Item Facility (IF): This is a measure of the difficulty of the item and is computed by dividing the total number of points actually earned on that item by a given subset of the test-taking population, by the total number of points that could have been earned had everyone gotten it right, and results in a decimal number between 0 and 1. For the all-or-nothing case in which partial credit is not awarded (e.g., the typical multiple choice test) this reduces to the number of people in the group getting the item right, divided by the total number of people in the group. Here, this was computed for the entire population, as well as for each of the high-, middle-, and low-scoring groups, each consisting of about one-third of the total population, as determined by ranking on the basis of the total test score, the latter being necessary in the following computation.together with plots of IF versus ID.2. Item Discrimination (ID): This reflects the correlation between getting the given item right and getting a high score on the test as a whole; i.e., the "predictive power" of the individual item. It is simply the difference of the IFs for the high- and low-scoring groups, and could in principle lie anywhere between -1 and 1, although a negative value would be extremely unlikely. (It could happen, say, if only high-scorers read something unintended into an, in fact, simple question that led to them arriving at the wrong answer while others got it right.)
The numerical statistics and graphs (1 marker = 1 test-taker) appear in Figures 1 to 5. The overall total test score profile is shown in Figure 1 and would appear to follow a basically normal distribution. High and low scores were 43.0 and 5.5, respectively, out of a possible 49 points, with a sample mean of 24.7 and standard deviation 8.3.
Figure 2 shows the IF values for the entire group adjacent to the ID values, followed by Figure 3 which shows the IFs for each of the high-, middle- and low-scoring thirds. Items with very high or very low overall IFs frequently have low discriminating power and should be classified as "poor" items. We can see that a fair number of items on this test fall into those categories, e.g., #2 and #33 (low IF) and #38 and #40 (high IF), which both have very low ID values. As would be expected, IF tends to increase from the low-scoring ("0") to the high-scoring ("2") group, although in some cases IF for the low group does exceed that of the middle group. On the whole, these also turn out not to be "good" questions.
Figure 4 shows IF versus ID on a grid for the 49 test items. Note the "inverted V" contour, with very easy or very difficult items of low discriminating power occurring near the lower right and left corners, respectively. Patterns showing only the right "leg" have also been observed in data sets obtained from other test administrations at my institution. (Note: Three items (#17, #31 and #34) obscured by other items on the original computer-generated plot have been replotted at adjacent locations here for the sake of clarity and completeness.)
Figure 5 offers a summary of the Figure 4 plot, in which the test items have been classified according to the scheme suggested in MacGregor [2]. In addition, the items of the present test have been grouped according to general type (answer the question, dictate the sentence, or translate into Japanese), and the question numbers have been suffixed to indicate the form of the question (WH-, yes/no, or choice). Most striking, perhaps, is the absence of pure dictation items (B: "Dictate the English sentence as is") from the top half of the chart, while 5 of the 7 translation items (C: "Translate the dictated English sentence into Japanese") rank in the "good" categories on the basis of ID value. It is not immediately apparent what differentiates the "good" answer-the-question items from the bad; it may be significant that more than half of the "poor" ones are numbered among items #25-#35, which related to study materials other than the "main" material for the class, although several top-row items also occur within this group.
Because of the limitations inherent in the data sets used thus far--in particular, the fact that no consideration had been given at the time the tests were prepared to the possibility of their ever being used in this type of test item analysis--it is risky to "conclude" anything at all. However, the fact that successful translation of dictated English sentences correlated very well with overall test scores may well merit further study, as would a more detailed look at the interplay (if it exists) between the grammatical form of a question and its ability to measure aspects of general linguistic ability and mastery of material content. The finding of appropriate visual representations for acquired quantitative data will certainly be prerequisite to the proper interpretation of that data as well.
Score range, frequency: [ 0.0, 2.0) < 0> | [ 2.0, 4.0) < 0> | [ 4.0, 6.0) < 1> |= [ 6.0, 8.0) < 0> | [ 8.0, 10.0) < 1> |= [ 10.0, 12.0) < 1> |= [ 12.0, 14.0) < 2> |== [ 14.0, 16.0) < 7> |======= [ 16.0, 18.0) < 6> |====== [ 18.0, 20.0) < 13> |============= [ 20.0, 22.0) < 5> |===== [ 22.0, 24.0) < 9> |========= [ 24.0, 26.0) < 9> |========= [ 26.0, 28.0) < 5> |===== [ 28.0, 30.0) < 7> |======= [ 30.0, 32.0) < 5> |===== [ 32.0, 34.0) < 6> |====== [ 34.0, 36.0) < 3> |=== [ 36.0, 38.0) < 2> |== [ 38.0, 40.0) < 3> |=== [ 40.0, 42.0) < 3> |=== [ 42.0, 43.0) < 2> |==
Item facility (IF) overall: Item, Item discrimination (ID): Item 1 => 0.778 |======= Item 1 => 0.364 |=== Item 2 => 0.156 |= Item 2 => 0.187 |= Item 3 => 0.339 |=== Item 3 => 0.457 |==== Item 4 => 0.311 |=== Item 4 => 0.377 |=== Item 5 => 0.544 |===== Item 5 => 0.488 |==== Item 6 => 0.194 |= Item 6 => 0.298 |== Item 7 => 0.711 |======= Item 7 => 0.374 |=== Item 8 => 0.289 |== Item 8 => 0.378 |=== Item 9 => 0.444 |==== Item 9 => 0.699 |====== Item 10 => 0.556 |===== Item 10 => 0.521 |===== Item 11 => 0.472 |==== Item 11 => 0.601 |====== Item 12 => 0.489 |==== Item 12 => 0.542 |===== Item 13 => 0.856 |======== Item 13 => 0.221 |== Item 14 => 0.489 |==== Item 14 => 0.560 |===== Item 15 => 0.539 |===== Item 15 => 0.628 |====== Item 16 => 0.428 |==== Item 16 => 0.356 |=== Item 17 => 0.378 |=== Item 17 => 0.572 |===== Item 18 => 0.456 |==== Item 18 => 0.647 |====== Item 19 => 0.644 |====== Item 19 => 0.646 |====== Item 20 => 0.767 |======= Item 20 => 0.384 |=== Item 21 => 0.283 |== Item 21 => 0.264 |== Item 22 => 0.383 |=== Item 22 => 0.570 |===== Item 23 => 0.506 |===== Item 23 => 0.783 |======= Item 24 => 0.156 |= Item 24 => 0.154 |= Item 25 => 0.422 |==== Item 25 => 0.256 |== Item 26 => 0.361 |=== Item 26 => 0.637 |====== Item 27 => 0.789 |======= Item 27 => 0.150 |= Item 28 => 0.728 |======= Item 28 => 0.175 |= Item 29 => 0.244 |== Item 29 => 0.108 |= Item 30 => 0.439 |==== Item 30 => 0.220 |== Item 31 => 0.483 |==== Item 31 => 0.542 |===== Item 32 => 0.528 |===== Item 32 => 0.653 |====== Item 33 => 0.078 | Item 33 => 0.194 |= Item 34 => 0.278 |== Item 34 => 0.380 |=== Item 35 => 0.278 |== Item 35 => 0.124 |= Item 36 => 0.817 |======== Item 36 => 0.293 |== Item 37 => 0.300 |=== Item 37 => 0.541 |===== Item 38 => 0.922 |========= Item 38 => 0.161 |= Item 39 => 0.639 |====== Item 39 => 0.480 |==== Item 40 => 0.983 |========= Item 40 => 0.036 | Item 41 => 0.656 |====== Item 41 => 0.329 |=== Item 42 => 0.661 |====== Item 42 => 0.238 |== Item 43 => 0.322 |=== Item 43 => 0.169 |= Item 44 => 0.917 |========= Item 44 => 0.109 |= Item 45 => 0.572 |===== Item 45 => 0.464 |==== Item 46 => 0.656 |====== Item 46 => 0.289 |== Item 47 => 0.294 |== Item 47 => 0.438 |==== Item 48 => 0.956 |========= Item 48 => 0.107 |= Item 49 => 0.178 |= Item 49 => 0.181 |=
Item Facility (IF) by score group (side by side): Group: -- 0 -- 1 -- 2 Item 1 => 0.571 0.806 0.935 |===== |======== |========= Item 2 => 0.071 0.129 0.258 | |= |== Item 3 => 0.107 0.323 0.565 |= |=== |===== Item 4 => 0.107 0.323 0.484 |= |=== |==== Item 5 => 0.286 0.548 0.774 |== |===== |======= Item 6 => 0.089 0.097 0.387 | | |=== Item 7 => 0.464 0.806 0.839 |==== |======== |======== Item 8 => 0.089 0.290 0.468 | |== |==== Item 9 => 0.107 0.387 0.806 |= |=== |======== Item 10 => 0.286 0.548 0.806 |== |===== |======== Item 11 => 0.125 0.532 0.726 |= |===== |======= Item 12 => 0.232 0.435 0.774 |== |==== |======= Item 13 => 0.714 0.903 0.935 |======= |========= |========= Item 14 => 0.214 0.452 0.774 |== |==== |======= Item 15 => 0.179 0.597 0.806 |= |===== |======== Item 16 => 0.321 0.274 0.677 |=== |== |====== Item 17 => 0.089 0.355 0.661 | |=== |====== Item 18 => 0.143 0.403 0.790 |= |==== |======= Item 19 => 0.321 0.613 0.968 |=== |====== |========= Item 20 => 0.536 0.823 0.919 |===== |======== |========= Item 21 => 0.107 0.355 0.371 |= |=== |=== Item 22 => 0.107 0.339 0.677 |= |=== |====== Item 23 => 0.071 0.548 0.855 | |===== |======== Item 24 => 0.071 0.161 0.226 | |= |== Item 25 => 0.357 0.290 0.613 |=== |== |====== Item 26 => 0.089 0.242 0.726 | |== |======= Item 27 => 0.786 0.645 0.935 |======= |====== |========= Item 28 => 0.696 0.613 0.871 |====== |====== |======== Item 29 => 0.214 0.194 0.323 |== |= |=== Item 30 => 0.393 0.306 0.613 |=== |=== |====== Item 31 => 0.232 0.419 0.774 |== |==== |======= Item 32 => 0.250 0.403 0.903 |== |==== |========= Item 33 => 0.000 0.032 0.194 | | |= Item 34 => 0.071 0.290 0.452 | |== |==== Item 35 => 0.214 0.274 0.339 |== |== |=== Item 36 => 0.643 0.855 0.935 |====== |======== |========= Item 37 => 0.071 0.194 0.613 | |= |====== Item 38 => 0.839 0.919 1.000 |======== |========= |========== Item 39 => 0.375 0.661 0.855 |=== |====== |======== Item 40 => 0.964 0.984 1.000 |========= |========= |========== Item 41 => 0.429 0.758 0.758 |==== |======= |======= Item 42 => 0.536 0.661 0.774 |===== |====== |======= Item 43 => 0.250 0.290 0.419 |== |== |==== Item 44 => 0.875 0.887 0.984 |======== |======== |========= Item 45 => 0.375 0.484 0.839 |=== |==== |======== Item 46 => 0.518 0.629 0.806 |===== |====== |======== Item 47 => 0.143 0.145 0.581 |= |= |===== Item 48 => 0.893 0.968 1.000 |======== |========= |========== Item 49 => 0.125 0.097 0.306 |= | |===
Item Facility (IF) vs Item Discrimination (ID) plot
IF: Left= 0.0 (hard), Right= 1.0 (easy)
ID: Top= 1.0 (good), Bottom= 0.0 (poor)
IF= 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
ID=1.0 ---------------------------------------------------------------------------------------------------------------
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
ID=0.9 ---------------------------------------------------------------------------------------------------------------
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
ID=0.8 ---------------------------------------------------------------------------------------------------------------
| | | | | |23 | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
ID=0.7 -------------------------------------------------09------------------------------------------------------------
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | 18 | 32 | 19 | | | |
| | | | 26 | | | | | | |
| | | | | | 15 | | | | |
| | | | | | | | | | |
ID=0.6 ----------------------------------------------------11---------------------------------------------------------
| | | | | | | | | | |
| | | | 2217 | | | | | |
| | | | | 14 | | | | |
| | | | | | | | | | |
| | | 37 | 3112 | | | | |
| | | | | | 10 | | | | |
| | | | | | | | | | |
ID=0.5 ---------------------------------------------------------------------------------------------------------------
| | | | | | 05 | | | | |
| | | | | | | 39 | | | |
| | | | 03 | | 45 | | | | |
| | | | | | | | | | |
| | | 47 | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
ID=0.4 ---------------------------------------------------------------------------------------------------------------
| | | | | | | | 20 | | |
| | | 340804 | | | |07 | | |
| | | | | | | | 01| | |
| | | | | 16 | | | | | |
| | | | | | | | | | |
| | | | | | | 41 | | | |
| | | | | | | | | | |
ID=0.3 ---------------------06----------------------------------------------------------------------------------------
| | | | | | | 46 | | 36 | |
| | | | | | | | | | |
| | | 21| | | | | | | |
| | | | | 25 | | | | | |
| | | | | | | 42 | | | |
| | | | | 30 | | | | 13 | |
| | | | | | | | | | |
ID=0.2 ---------------------------------------------------------------------------------------------------------------
| 33| 02 49| | | | | | | | |
| | | | 43 | | | | 28 | | |
| | | | | | | | | | 38 |
| | 24 | | | | | | 27 | |
| | | | | | | | | | |
| | | 35| | | | | | | |
| | | 29 | | | | | | | 44 48 |
ID=0.1 ---------------------------------------------------------------------------------------------------------------
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | |
| | | | | | | | | | 40|
| | | | | | | | | | |
| | | | | | | | | | |
ID=0.0 ---------------------------------------------------------------------------------------------------------------
General classification of test items according to Item Facility (IF),
Item Discrimination (ID) and question type (A,B,C; see below):
Difficulty? IF < 0.3 0.3 <=IF<= 0.7 IF > 0.7
(Too hard) (Reasonable) (Too easy)
Discrimination?
--------------------------------------------------------------
| A | A #05 ,#09 ,#10 | A |
| | #11 ,#12 ,#14y | |
ID > 0.4 | | #15 ,#17 ,#18 | |
(Good) | | #19c,#22 ,#23 | |
| | #26 ,#31 ,#32 | |
| | | |
| B | B | B |
| | | |
| C #47 #37 C #39 ,#45 | C |
--------------------------------------------------------------
| A #08 ,#34 | A #04 ,#16c | A #01 ,#07c, #20 |
0.3<=ID<0.4 | | | |
(Good, but | | | |
might be | | | |
improved) | B | B | B |
| | | |
| C | C #41 | C |
--------------------------------------------------------------
| A #06 ,#21 | A #25 ,#30 | A #13c |
| | | |
0.2<=ID<0.3 | | | |
(Marginal) | | | |
| B | B #42 ,#46 | B #36 |
| | | |
| C | C | C |
--------------------------------------------------------------
| A #02y,#24 ,#29 | A #16 | A #27 ,#28 |
| #33 ,#35 | | |
ID < 0.2 | | | |
(Poor) | B | B | B #38 ,#40 ,#44 |
| | | #48 |
| | | |
| C #49 | C #43 | C |
--------------------------------------------------------------
Question type: A = Answer the English question in English
[no letter: WH-question
y : yes-no question
c : choice question]
B = Dictate the English sentence as is
C = Translate the dictated English sentence into Japanese
The author, Bill Rockenbach, can be reached by email at:
and welcomes suggestions and comments concerning this material, as well as
inquiries regarding any data sets referenced herein. For other information,
please visit his website at:
Please note, however, that he is based in Japan, so communications to and
from the U.S. will generally be delayed by at least half a day.
This file last modified: Thursday, 14-Sep-2000 00:47:16 JST Japan Standard Time You accessed it: Wednesday, 19-Aug-2026 20:16:13 JST Japan Standard Time