Several quantities have been proposed for quantitative assessment of glaucoma. The present study intended to estimate critical limits for not pathological eyes that can be used to classify eyes as glaucomatous in the population sampled. Such critical limits are necessary for estimation of the relative sensitivity for detection of glaucoma for quantities estimated in the UGDS study. The limited sample size is associated with uncertainty and was considered by expressing the critical limit as an extreme confidence limit. Critical limits intended for clinical differentiation between glaucoma and not glaucoma should be derived from a sample size of at least 250.
In the present study, the non-glaucoma population sampled consisted of eyes clinically not considered glaucoma suspect in subjects where the fellow eye was clinically considered glaucoma suspect. Therefore, the critical limit estimated for differentiation between glaucoma and not glaucoma is valid only for initial clinical evaluation of eyes that may later convert to glaucoma.
The different quantities presently analyzed were often but not always measured in the same eye. All estimates are therefore dependent. The comparison of estimates is consequently limited to description. However, since all measurements for all quantities are related to the same cohort of subjects it is anticipated that discrepancies between the quantities express real differences.
The variation of sample size for different quantities was caused by variable subject cooperation, and occasional device malfunction requiring repair and re-calibration by the distributor (Table 5, footnote 1). The variation impacts the precision in estimates of mean and standard deviation and indirectly impacts the precision expressed in the confidence intervals for the variation coefficients and the tolerance limits.
The cpRNFLT-Global measurements were derived from 3 different generations of OCT devices representing technology steps that allow increasingly faster sampling. Increased sampling rate translates into increasing measurement precision since eye movement limits the precision within 1 ONH capture. In the current paper, analysis was made on the cpRNFLT-Global data collected with the 3 different OCT-devices pooled together and separated into data captured with time-domain and frequency domain, respectively (Table 6).
The scatterplots demonstrated that the presently analyzed sub-cohort of not glaucoma suspect eyes covered the age interval when glaucoma typically occurs and that both sexes were evenly distributed over the age interval (Fig. 1).
The lack of iterated measurements of IOP is a limitation of the data available. For this reason, the value of averaging over several IOP measurements at each occasion cannot be conclusively judged. The variance for subjects and occasions was of the same order (Table 5) and the variance for occasions estimated with the model in Appendix 2, Eq. 2, is the sum of variances for occasions and iterations. It is probable that iteration of IOP measurements would improve precision.
For the MD estimates, the estimated variance for subjects (Table 5) expresses the subject variability around the reference population used in the Humphrey machine since MD is age and sex adjusted to the reference population.
The sum of the variance for subjects and occasions was large for all the quantities analyzed (Table 5). At the first visit, a measurement necessarily has to be compared to a critical limit defined by the variance for subjects. Since the variance for subjects is substantial (Table 5) such an evaluation is inefficient.
For IOP, the precision can probably be increased by averaging over iterations. Also, for MD, precision would probably improve by averaging over several measurements at each occasion. However, this would be very cumbersome in clinical routine work due to the time consumption and generate unrealistic burden for the patient. For the other quantities averaging over three iterations would be possible in clinical routine work and is desirable to gain some precision (Appendix 3, Eq. 4).
Glaucoma is supposed to cause either an increase or a decrease depending on the quantity measured. Therefore, one sided tolerance limits were estimated. On the condition that the normal distribution is a good approximation, critical limits defined as tolerance limits may be defined based on a known expected population mean and standard deviation, respectively. However, in the current study the expected mean and standard deviation were estimated from a limited sample. Therefore, extreme confidence limits for one sided tolerance limits were estimated [14].
Typically, critical limits should be calibrated to age and sex. Our finding that the MD measurements are independent of age and sex is consistent with the fact that MD by definition is age and sex adjusted. Other studies have shown that cpRNFLT-Global decreases with age [1], [10], [11]. Considering the variability of cpRNFLT-Global among subjects the sample size in the current study was too small to resolve a cpRNFLT-Global loss with age (Fig. 1, Table 5). For the other quantities measured, it is similarly possible that the substantial variability among subjects (Table 5) obscure a slight age and or sex dependence.
Critical limit estimates based on the normal distribution requires data that are approximately normal distributed. The deviation from the normal distribution at extreme values in the NED plots (Fig. 2) is expected due to the limited sample size. Statistical verification of an underlying normal distribution would require a much larger sample than was available. The observations for all the quantities measured within 95% around the mean are close to normal distributed (Fig. 2). The comparison of critical limit distance from mean normalized to mean among quantities measured (Table 6, last column) is therefore relevant. The C/D-linear measurements are relative numbers between 0 and 1. The frequency distribution is therefore expected to be truncated close to the extreme values. However, the estimated mean C/D-linear was centered at 0.6 (Table 6). Variation around the center between the extreme limits is therefore expected to be approximately normal distributed. This was supported by the NED-plot for C/D-linear measurements (Fig. 2).
The precision of measurement of eye within subjects depends on variability among subjects, occasions within subjects, and measurements (Appendix 5). The analysis of variance demonstrated that the variance for subjects dominated for all quantities estimated (Table 5). This agrees with previous findings for NRA [32]. Consequently, the capacity to distinguish pathological from not pathological is, for all critical limits derived, limited by substantial variability among subjects. This implies that critical limits derived from population estimates of mean and standard deviation are insensitive for detection of pathological. Sensitivity would increase substantially if critical limits derived from within subjects measurements, i.e. multiple occasions, are applied.
The currently observed variabilities for IOP measurements (Table 5) were of the same order that have previously been published [8], 29. Several IOP measurements could easily be iterated in clinical routine work. Averaging would reduce the total variation for subjects (Eq. 4). The presently observed variabilities among MD measurements are comparable to previously published data (Russell, Garway-Heath & Crabb 2013). Averaging of several measurements at the same occasion is clinically impossible due to time consumption of the measurement procedure and inconvenience for the patient. The variability among subjects was larger for Stratus than Cirrus (Table 5). Previous estimates of precision of subject measurements are of the same order and found a similar difference [20], [27], [34]. The presently observed lower variabilities for cpRNFLT-Global measurements with the Cirrus device than for the Stratus device (Table 5) translates into shorter critical limit distance from the mean, normalized to the mean (Table 6). The here observed variabilities for the morphometric variables were approximately similar to previously published data: C/D-linear, NRA-Global, cpRNFLT-Global, GDx-TSNIT [9], [12], [25]. For the morphometric quantities measured, the variance for iterations was low in relation to the sum of the variance for subjects and occasions (Table 5). Therefore, averaging over several iterations at each occasion is expected to have only limited impact on measurements in subjects (Eq. 4).
In the current study, the estimated extreme confidence limit for the one sided 95% tolerance limit corresponds to a specificity of 95 % for the quantity indicated. The currently estimated critical limits for IOP, C/D-linear, and NRA-Global correspond to the previously published critical limits (Table 2). The lower-than-expected critical limits for GDx-TSNIT and cpRNFLT may arise from device-specific precision and calibration and be associated with differences in populations measured. These differences emphasize the importance of establishing clinical center based critical limits. A small sample size when estimating the confidence limit for the tolerance limit introduces uncertainty. Thus, a confidence limit for a small sample is further away from the mean than for a large sample. The small sample size available for GDx-TSNIT therefore contributes to the low estimated critical limit (Table 6, 6th column) in comparison to previous data (Table 2, last row).
Not considering MD, the critical limits presently established, suggest that for glaucoma detection a larger relative change from expected mean in not glaucoma suspect eye is required for the HRT quantities measured than for the other quantities measured (Table 6, last column). This is consistent with previous estimations of variability [32].
It should be pointed out that a critical limit based on level and variability of a measured quantity defines the probability to wrongly classify a measurement as pathological although it is not (1 minus the tolerance selected, where tolerance is equivalent to specificity). The clinical significance of the critical limit depends on the importance of a defined change with regard to the disease, in the present study glaucoma. If an IOP above the statistically defined critical limit does not impact the progress of the disease, measurements of IOP are not suitable for detecting glaucoma even if the critical limit is close to the expected mean in non-glaucoma eyes.
Critical limits for discrimination of pathology from non-pathology used at the initial patient visit necessarily includes variability among subjects and therefore are insensitive for detection of glaucoma. The current analysis indicates similar but better potential capacity to discriminate glaucoma from not glaucoma at any arbitrarily defined specificity for IOP, cpRNFLT-Global and GDx-TSNIT than for HRT measurements at the initial patient visit (Table 6, last column).
To evaluate the capacity of a quantity to correctly identify glaucoma, the phenotype of glaucoma in focus has to be uniquely specified. Further, it must be demonstrated that the expected mean of the quantity analyzed is related to the grade of the phenotype of glaucoma in focus. Thus, any estimation of sensitivity for a specific phenotype of glaucoma depends on the definition of the phenotype.
It is concluded that it is preferable to estimate critical limit for small samples as extreme confidence limit for the tolerance limit. The presently estimated critical limits are comparable to previous findings for the device models used. C/D-linear, and NRA-Global estimated with HRT provides lower efficiency for estimating sensitivity than the other quantities measured. Specific intra-individual critical limits require a small increase of the measured quantity to identify glaucoma when the patient has glaucoma.
Comments (0)