How JPEG Evaluates Subjective Image Quality
Evaluating subjective image quality across international laboratories requires the Joint Photographic Experts Group (JPEG committee) to implement rigorous, standardized assessment protocols. This process relies on internationally harmonized testing methodologies, calibrated viewing environments, diverse test datasets, and advanced statistical cross-verification. By coordinating these efforts across research institutions worldwide, the JPEG committee ensures that image compression standards are judged reliably against human visual perception rather than algorithmic approximations alone.
Standardized Assessment Methodologies
The JPEG committee primarily adheres to assessment frameworks defined by the International Telecommunication Union, notably ITU-R BT.500 and ITU-T P.910, alongside dedicated standards such as ISO/IEC 29170-2. Laboratories employ specific psychophysical test paradigms depending on the evaluation target:
- Double Stimulus Impairment Scale (DSIS): Observers view an uncompressed original reference image followed by the compressed target image, rating the degree of perceived impairment on a discrete five-point scale.
- Double Stimulus Continuous Quality Scale (DSCQS): Observers evaluate pairs of images (reference and coded) in random order without knowing which is the original, providing ratings on a continuous scale to detect subtle compression artifacts.
- Simultaneous Side-by-Side Comparison: Often applied in flicker tests or side-by-side preference ratings, allowing viewers to inspect specific spatial regions under identical temporal conditions.
Environmental and Display Calibration
To eliminate hardware-induced variability across different global sites, testing laboratories must calibrate their physical environments to uniform specifications:
- Display Settings: Monitors are calibrated to target luminance levels (typically 100 to 200 cd/m² for standard dynamic range, or standard peak luminance for HDR), specific color gamuts (such as sRGB, DCI-P3, or BT.2020), and standard white points (usually D65).
- Viewing Distance: Distances are locked proportionally to the monitor’s height or pixel density (commonly 3H to 4H, where H is display height) to standardize visual angle and angular resolution in cycles per degree.
- Room Illumination: Ambient lighting is maintained at low, controlled levels (typically under 20 lux) with neutral gray room surroundings to prevent glare and maintain consistent contrast perception across labs.
Source Dataset and Test Material Curation
Test imagery is curated to represent a comprehensive cross-section of visual challenges. The selected dataset includes varying spatial frequencies, complex textures, smooth gradients, diverse skin tones, text, and differing dynamic ranges. Standardized software distributes identical source files without applying any local scaling, operating system color correction, or unauthorized pre-processing.
Observer Selection and Screening Protocols
Subjective evaluations require human panels comprising non-expert observers to reflect general consumer perception. Each laboratory applies strict pre-test screening:
- Visual Acuity: Assessed using standard Snellen or Landolt C charts to ensure normal or corrected-to-normal vision.
- Color Vision: Checked via Ishihara plates or the Farnsworth-Munsell test to exclude color-deficient participants.
- Sample Size: Each testing facility typically recruits at least 15 to 20 valid observers per experiment to satisfy statistical power requirements.
Cross-Laboratory Data Harmonization and Analysis
Once testing concludes, raw data from all contributing international laboratories undergo centralized statistical processing to verify reproducibility:
- Subject Screening and Outlier Removal: Standardized algorithms (as outlined in ITU-R BT.500) identify and eliminate participants whose ratings show excessive internal inconsistency or deviate significantly from the group consensus.
- Score Normalization: Individual ratings are converted into Mean Opinion Scores (MOS) or Differential Mean Opinion Scores (DMOS). Z-score normalization is applied to compensate for cultural differences in how observers across regions utilize numerical rating scales.
- Cross-Lab Correlation: Pearson linear correlation coefficients (PLCC) and Spearman rank-order correlation coefficients (SROCC) are calculated between laboratories. A high correlation confirms that the observed visual performance is standard-compliant, statistically robust, and independent of geographic or laboratory bias.