| Jiwon Park | 2 Articles |
Purpose
This study evaluates the performance of Claude and GPT LLM Vision APIs for automated clinical questionnaire processing in spine surgery by comparing accuracy, efficiency, reproducibility, and cost-effectiveness. Methods Clinical questionnaires from 56 patients (336 total pages) were processed using a Python 3.12-based system incorporating PDF preprocessing, image enhancement via OpenCV, and direct LLM Vision analysis. Both models were evaluated on 26 questionnaire items (1,456 data points) using accuracy comparison, processing time measurement, token utilization analysis, and intra-class correlation coefficient (ICC) assessment through three independent iterations. Results GPT achieved 98.83% accuracy (1,439/1,456) compared to Claude's 97.94% (1,426/1,456). Both models processed questionnaires in 27 seconds per set, representing 68% time reduction versus manual entry (85 seconds). GPT demonstrated 59% cost advantage ($0.023 vs. $0.056 per questionnaire), while Claude showed superior reproducibility (ICC 0.98 vs. 0.96). GPT achieved 100% accuracy across 21 items versus Claude's 17 items. Error analysis identified predominantly handwriting recognition (52%) and image quality issues (28%), with 89% of errors successfully flagged for review. Conclusions Both models achieve clinical-grade performance exceeding 90% accuracy. GPT demonstrates superior accuracy and cost-effectiveness, while Claude provides better reproducibility. Model selection should be guided by institutional priorities regarding accuracy, reproducibility, and operational scale.
Purpose
This study aimed to evaluate the predictive performance of preoperative computed tomography–based Hounsfield unit (HU) and magnetic resonance imaging-based vertebral bone quality score (VBQS) for cage subsidence after 1- to 2-level oblique lumbar interbody fusion (OLIF), and whether it differs by intraoperative cage placement position. Methods Ninety-one OLIF levels in 54 patients (2015–2022) were retrospectively reviewed in this single-center cohort. Subsidence was defined as ≥2 mm middle disc-height reduction or ≥2 mm cage protrusion at 1 year. Multivariable logistic regression with cluster-robust standard errors estimated adjusted odds ratios (OR); the pre-specified primary test was the lower-instrumented-vertebra HU×cage-position interaction. Results Subsidence occurred in 24 of 91 levels (26.4%). Each one-standard-deviation decrease in lower-instrumented-vertebra HU was independently associated with subsidence (adjusted OR, 0.40; 95% confidence interval, 0.20 to 0.81; p=0.011); VBQS was not. The subsidence group included more osteoporosis-range levels (<110 HU; 62.5% vs. 25.4%, p=0.003). The interaction was not significant (p=0.55), but the HU effect concentrated in middle-placed cages (adjusted OR, 0.345; p=0.012) and attenuated in anterior-placed cages (OR, 0.50; p=0.18). Conclusion Lower-instrumented-vertebra HU is an independent predictor of cage subsidence after 1- to 2-level OLIF, most evident in middle-placed cages.
|
|