EducationPsychology

Inter-Judge Reliability of the Indiana State School Music Association High School Instrumental Festival

Timothy D. Brakel

2006.10.1JOURNAL OF BAND RESEARCH

Abstract

Abstract The study examined issues of reliability as related to the Indiana State School Music Association Instrumental Festival. Using the total 2002 and 2003 population of ISSMA judges' panels (n=43 panels consisting of 3 judges per panel) and events (n=840), inter-judge reliabilities were computed for each panel regarding overall reliability, agreement between pairs of judges, type of organization (band, string orchestra or full orchestra) adjudicated and the group level that the ensemble performed in. Each ensemble was rated in the categories of intonation, tone quality, articulation/tonguing/bowing techniques, technique/fluency/mechanical skill, note accuracy, interpretation/musicianship, dynamics, balance and blend, and other factors. Each category was rated on a scale of one (best) to four (worst). The criterion for reliability was the final sum point total of all categories. The reliabilities were found to be overall acceptable for band and orchestra adjudication. The judges training session prior to the 2003 festival improved inter-judge reliability-especially for orchestral ensembles. Reliabilities for 3-member panels were better than for pairs of judges. Adjudication of poor performances appeared to result in greater inconsistency between judges. Instrumental Festival The reliability of judges' contest ratings is of great interest to music directors. Music programs and their directors are in need of valid and reliable assessment models. In several studies, the reliability for the overall festival rating appears to indicate good reliability while ratings for individual captions were generally found to be less reliable (Burnsed, Hinkle and King, 1985; Burnsed and King, 1987; Decamp, 1980; Garman, Boyle & DeCarbo, 1991; Nichols, 1985). This may be due to judges who predetermine the overall rating and then adjust the individual captions' scores to achieve the desired rating. In most, if not all, contest situations, the captions for the evaluation form were selected on the basis of face validity. Little research has been conducted as to whether the categories are appropriate. Few organizations that sponsor contests conduct research regarding the reliability of the rating form, formal evaluation of judges, and other issues related to performance. That is not to say that these organizations are not interested in issues related to reliability since many of these organizations sponsor judges' training sessions with the goal of improving judge reliability. Recordings of performances for evaluation by a panel of judges have been suggested as a means of improving reliability (Fiske, 1983; Massell, 1978; Vasil, 1973). This places great responsibility on the recording technology and the recording engineer to produce a high quality recording. The advantage of recordings is that each judge would hear the same performance from the same vantage point- i.e., the placement of the microphones. One disadvantage is that recordings lose some degree of immediate feedback to the performing organization since the performance would first have to be recorded and then adjudicated at a later time. Forbes (1994) suggested that one way to improve judge reliability was to employ judges who demonstrate high degrees of reliability. Forbes believed that such a procedure would be difficult to initiate but would improve the reliability of adjudicators in the long term. Some researchers suggest that a panel of five to twelve judges improves reliability (Fiske, 1975; Vasil, 1973). Bergee (2003) found that university faculty member jury panels exhibited greater variability when jury panels of two or three were used than for panels consisting of four or five faculty. Employing panels of five or more judges would potentially add considerable expense to the organization sponsoring the event. It also may lead to less objective measurement since many contest situations use tape recorded comments about the performance; having several judges in close proximity to each other may create a situation where one judge may overhear another's comments. …

Citation format

BRAKEL, Timothy D. Inter-judge reliability of the indiana state school music association high school instrumental festival. JOURNAL OF BAND RESEARCH, 2006, 42: 59.