JEFF CUBOS
  • Blog
  • Reviews
    • CE Reviews
    • Research Reviews
    • Book Notes
  • About

Gait Parameters & Tibial Stress-Fracture Risk - A Conversation with Claude

8/10/2026

0 Comments

 
This one is a little bit different. 

I came across this paper and immediately upon scanning I realized that this would be one of those studies that would be way over my head. One I normally give up on and don’t end up reading in its entirety..,if at all.

Simply put, I don’t do well with models and modelling. 

So I asked Claude for some help.

Not to read the paper for me. But to better help me understand both the nuances and the relevance of it.

It was an interesting experience. A beneficial one actually.

You’ll see my questions below. I’m sure I could have asked a lot more. Next time, I’m sure I will.

But at the end of the day, this conversation with an LLM helped me work through the parts I didn’t understand, in order to fully appreciate and question what their findings actually meant. For me. In the clinic.

So below is the review in transcript form of my conversation with Claude:
​
*Note: Reading the article FOLLOWING this conversation was much easier than it would have been had I not met Claude for coffee.
​
Picture
Nagaraja S, Rubio JE, Subramaniyan M, Tong J, Baggaley M, Edwards WB and Reifman J (2026) Identification of modi able gait parameters predictive of simulated tibial stress-fracture risk in men and women. Front. Physiol. 17:1905928. doi: 10.3389/fphys.2026.1905928

Gait Parameters & Tibial Stress-Fracture Risk — Conversation with Claude

[Uploaded PDF: "Identification of modifiable gait parameters predictive of simulated tibial stress-fracture risk in men and women" — Nagaraja et al., Frontiers in Physiology, 2026]

Summarize this paper for me to make it easier to comprehend

C

Here's a plain-language summary of the paper:

What they did Researchers wanted a faster, field-usable way to predict which people are at risk of tibial stress fractures (a common overuse injury in military training) based on their running gait. They used data from 41 healthy young adults (21 men, 20 women) who ran with three different loads (0, 11.3, or 22.7 kg) on a treadmill. Previous work had already computed each person's "true" stress-fracture risk using detailed, computationally expensive biomechanical models. This study built simpler statistical models to approximate those results using 30 gait measurements (like joint angles, forces, and moments).

Two types of models were built:

  1. Risk prediction models – estimate someone's actual percentage risk
  2. Risk classification models – simply flag someone as "low" or "high" risk

Key findings

  • Ankle joint reaction force (basically, force through the ankle) was the single strongest predictor of fracture risk for both men and women.
  • The next most useful factor differed by sex: hip external rotation angle for men, hip adduction moment for women.
  • Using these top predictors, the classification models were quite accurate (AUC 0.81–0.95, where 1.0 is perfect).
  • If they used only joint-angle data (which could theoretically be captured by wearable sensors in the field, without needing force-plate/lab equipment), accuracy dropped — especially for men (0.53–0.71, fairly weak) but stayed reasonable for women (0.72–0.82).
  • Ankle dorsiflexion angle was the best "angle-only" predictor for both sexes.

Why it matters This suggests that, in theory, simple wearable sensors measuring joint angles could eventually screen at-risk women reasonably well without expensive lab equipment, though this works less well for men, who need force data too. This could someday help identify military recruits or runners who need individualized training adjustments to prevent stress fractures.

Caveats the authors flag

  • Small sample size (41 people total)
  • Risk estimates are based on a computational model, not actual clinical fracture diagnoses
  • Findings need validation on new, independent data before real-world use

What are the paper's weaknesses

C

Based on the paper's own discussion and a critical read, here are the main weaknesses:

Acknowledged by the authors

  1. No real-world validation — The "ground truth" isn't actual diagnosed fractures; it's the output of an earlier computational model (musculoskeletal → finite-element → probabilistic model). The whole study is essentially predicting a model's prediction, not clinical outcomes. If that underlying model is biased, everything built on it inherits the bias.

  2. Very small sample size — Only 21 men and 20 women. This forced them to cap models at just 5 predictors to avoid overfitting, and makes it hard to trust that the "most informative" parameters would generalize to other populations.

  3. Circular/optimistic bias in the classification models — They used the same full dataset to first identify the top predictive parameters (via the LME models) and then built the logistic regression models using those same parameters. The authors admit this likely makes the reported AUC values optimistically inflated.

  4. No true independent test set — While they did 100 repetitions of 5-fold cross-validation, this is all still internal resampling of the same 41 people. There's no separate cohort to confirm the models actually generalize.

  5. Statistical non-independence — Each participant contributed 3 data points (one per load condition), but the AUC confidence intervals were calculated at the observation level, treating these as independent. This likely makes the confidence intervals falsely narrow (overconfident).

  6. Arbitrary risk thresholds — There's no clinically established cutoff for "low" vs. "high" fracture risk, so they tested arbitrary thresholds (2–8%) somewhat blindly, and cutoffs above 8% couldn't even be used due to severe class imbalance.

  7. Weak performance for men with kinematics-only models — The whole "field-deployable" promise (using just wearable sensors) largely falls apart for men, where AUCs dropped to 0.53–0.71 — barely better than chance in some cases.

Additional critiques

  1. Simulated combat training, not real training — The 10-week BCT "exposure" was simulated computationally, not observed in actual recruits going through real training with all its variability (terrain, fatigue accumulation, footwear, etc.).

  2. Narrow population — Participants were young (18–25), healthy, experienced recreational runners recruited via convenience sampling at one university — not necessarily representative of the broader, more diverse military recruit population (varying fitness levels, prior injury history, running experience).

  3. Treadmill-only, single speed — Running was tested at a fixed 3.0 m/s on a treadmill, which doesn't capture overground running variability, terrain changes, or varying paces typical of actual training.

  4. Exploratory, hypothesis-generating nature — The authors themselves label it "exploratory," and some of the identified associations (e.g., hip flexion angle in women, hip external rotation in men) have no prior literature support — meaning they could be dataset-specific artifacts rather than real biomechanical relationships.

In short: the methodology is sound and transparent, but the findings rest on a small sample, a computational (not clinical) outcome measure, and some methodological choices that likely inflate performance metrics — all of which the authors are commendably upfront about, but which mean this is very much a "proof of concept" rather than a validated clinical tool.


So how can a rehab clinician benefit from this paper

C

Here's how a rehab clinician (PT, athletic trainer, sports medicine provider) could realistically use this paper — with appropriate caveats:

1. Gait parameters to prioritize in assessment If working with runners or military trainees, this gives a shortlist of what to watch for during gait analysis:

  • Ankle joint reaction force — the single strongest predictor in both sexes. High ankle loading during running is a red flag.
  • Ankle dorsiflexion angle — the most useful observable kinematic parameter (measurable without force plates), for both sexes.
  • Hip external rotation angle — worth watching in men.
  • Hip adduction moment / hip flexion angle — worth watching in women.

This isn't new territory (prior literature already flagged ankle loading and hip mechanics), but this paper adds some sex-specific nuance clinicians can factor into individualized assessments.

2. Supports the case for sex-specific programming The consistent finding that men and women have different predictive gait signatures reinforces that generic "one-size-fits-all" injury-prevention gait retraining protocols may be suboptimal. A clinician could use this as justification for tailoring cueing and correction strategies differently by sex — e.g., more attention to hip external rotation control in male runners, more attention to hip adduction/frontal-plane hip control in female runners.

3. Realistic expectations about wearable/IMU-based screening If a clinic is considering IMU-based gait screening tools (increasingly common in sports medicine), this paper is a useful reality check: kinematic-only (angle-only) data works reasonably for flagging at-risk women but performed poorly for men, because kinetic (force/moment) data still matters and can't yet be fully substituted by wearables. This tempers overselling of "smart insole" or wearable screening products, especially for male patients.

4. A screening mindset, not a diagnostic one A clinician could take the general framework — track a few high-yield modifiable gait variables during load-carriage or return-to-run testing — as a lightweight monitoring tool for at-risk populations (e.g., new recruits, high-mileage runners returning post-injury), understanding it's a rough risk-stratification idea, not a validated diagnostic algorithm.

Important caution to keep in mind This is explicitly an early-stage, computationally-derived, small-sample study — not validated against real clinical fracture outcomes. A clinician shouldn't treat the specific AUC numbers or cutoffs as clinically actionable thresholds. It's best used as a way to prioritize what to look at during gait assessment and to justify why individualized, sex-specific approaches make biomechanical sense — not as a ready-made risk calculator.


How can a rehab clinician test for each of the above (in #1)

C

Here's how a clinician could realistically assess each parameter, ranging from simple clinical tests to more advanced tools:

1. Ankle joint reaction force (hardest to measure directly) True joint reaction force requires force plates + musculoskeletal modeling software — not available in most clinics. Practical proxies:

  • Force plate or instrumented treadmill (if available) — measure vertical ground reaction force (GRF) peak and loading rate, which correlates with ankle loading.
  • Pressure-sensing insoles / wearable force sensors (e.g., Loadsol, similar products) — increasingly affordable and can estimate loading during overground or treadmill running.
  • Impact sound/audio cue test — informal but useful: louder footstrike often correlates with higher impact forces.
  • Vertical oscillation and cadence — indirectly, low cadence and high vertical displacement are associated with higher impact loading; easy to assess visually or with a metronome + video.
  • Tibial acceleration — a wearable accelerometer strapped to the shin (research-grade, but some consumer options exist) gives a reasonable proxy for tibial loading.

2. Ankle dorsiflexion angle (during running — more feasible)

  • 2D video gait analysis — record sagittal-plane video at 120+ fps (most smartphones can do this) during treadmill running, then use frame-by-frame angle measurement (free tools like Kinovea) at initial contact and midstance.
  • Wearable IMUs on the shank and foot — commercial systems (e.g., Xsens, DorsaVi, IMeasureU) can quantify dynamic ankle angles during running, as referenced in the paper.
  • Static dorsiflexion screen as a proxy — weight-bearing lunge test (knee-to-wall) to assess available ankle dorsiflexion range; limited static ROM often predicts reduced dynamic dorsiflexion during gait.

3. Hip external rotation angle (men)

  • 3D motion capture (gold standard, but usually only in specialized labs/universities).
  • 2D frontal/transverse plane video — harder to capture accurately than sagittal-plane motion since rotation is axial, but qualitative assessment of foot progression angle during running can serve as a rough surrogate (excessive toe-out often reflects hip external rotation).
  • IMU-based systems with multiple sensors (pelvis + thigh) can estimate transverse-plane hip rotation dynamically.
  • Clinical exam correlate — passive hip internal/external rotation ROM in prone or seated position; reduced hip internal rotation ROM often associates with compensatory external rotation during gait.

4. Hip adduction moment (women)

  • Requires kinetic data (force + kinematics), so a true moment calculation needs 3D motion capture + force plates — typically research/hospital gait labs only.
  • Practical clinical surrogate: dynamic knee valgus / hip drop assessment
  • Single-leg squat test — watch for hip adduction/pelvic drop (Trendelenburg-type pattern), a widely used clinical proxy for hip adduction moment tendencies.
  • Step-down test — similar principle, assessing frontal-plane hip and knee control.
  • 2D video during running (rear view) — look for excessive hip adduction ("crossover gait," where the foot lands near or across the midline) — a well-established visual marker correlated with hip adduction moment.

5. Hip flexion angle (women)

  • 2D sagittal video analysis — measure hip flexion angle at initial contact/midstance using apps like Kinovea, Dartfish, or even Coach's Eye.
  • IMU-based wearables on thigh/pelvis for dynamic angle tracking.
  • Clinical correlate: assess hip flexor length/tightness (Thomas test) and active hip flexion ROM, since restricted mobility can alter dynamic gait mechanics.

Practical takeaway for a typical clinic (no motion-capture lab): A reasonable field-based protocol would combine:

  • Slow-motion smartphone video (sagittal + posterior views) during treadmill running
  • A single-leg squat or step-down test to screen frontal-plane hip control (adduction moment proxy)
  • A weight-bearing lunge test for ankle dorsiflexion mobility
  • Cadence/vertical oscillation as loading proxies
  • Referral to a lab or use of research-grade wearables (IMUs, instrumented insoles) if a more precise kinetic/kinematic profile is needed — especially useful given the paper's finding that kinetic data adds real value beyond angles alone, particularly for men

Ok to add to your practical takeaways, what about the following? Ankle joint reaction force: single leg hop/jump test or other?

Ankle ROM: can we use a goniometer?

Hip rotation angle: goniometer or table top?

Hip adduction moment: hand held dynamometer? Or visual closed kinematic chain test?

Hip flexion angle: goniometer or table top?

C

Good instincts — these are all legitimate clinical tools, but it's worth being precise about what they actually capture versus what the paper measured (dynamic, running-specific angles/forces). Here's how each holds up:

Ankle joint reaction force → single-leg hop/jump test? Reasonable as an indirect proxy, not a direct measure.

  • A single-leg hop test won't quantify force in newtons, but it can reveal loading strategy and control — landing mechanics, ability to absorb load softly vs. stiffly, asymmetries between limbs.
  • Single-leg hop for distance/time, drop vertical jump, or single-leg landing tests can flag poor eccentric control or stiff landing patterns that are plausible surrogates for elevated impact loading.
  • Better if paired with video analysis (loud/hard landing, minimal knee/hip flexion on landing = higher impact forces) or a force plate/force mat if available.
  • Bottom line: good functional screen for loading strategy, not a direct measure of ankle JRF magnitude — treat it as a qualitative red flag generator.

Ankle dorsiflexion ROM → goniometer? Yes, absolutely — this is the standard clinical tool.

  • Weight-bearing lunge test with goniometer (measuring knee-to-wall distance or the tibial angle) is well-validated and clinically practical.
  • Non-weight-bearing goniometric ankle dorsiflexion is also fine but weight-bearing is more functionally relevant to running mechanics.
  • Just remember: this gives you available static ROM, not dynamic dorsiflexion angle actually used during running — the paper's parameter is the latter. Limited static ROM is a reasonable predictor that dynamic ROM will also be limited, but it's not a 1:1 substitute.

Hip external rotation angle → goniometer or tabletop (prone) test? Yes, either works, and both are standard.

  • Prone hip rotation test with goniometer (knee bent to 90°, measuring internal/external rotation) is the classic exam.
  • Tabletop/seated hip rotation test is also valid and sometimes more comfortable for patients.
  • Same caveat as above: this measures passive/static rotational ROM, which is a reasonable clinical correlate but not identical to the dynamic transverse-plane hip rotation angle occurring mid-stance during running. Still clinically useful as a proxy and is far more accessible than lab-based rotational kinematics.

Hip adduction moment → hand-held dynamometer or visual closed-chain test? Both are useful, but they capture different things — worth using together.

  • Closed-chain visual tests (single-leg squat, step-down, Trendelenburg) are the more direct functional proxy, since hip adduction moment during running reflects dynamic frontal-plane control under load — closed-chain tests replicate that loading condition and are widely used clinically for this exact purpose.
  • Hand-held dynamometer (HHD) measuring hip abductor strength (open-chain, side-lying) is also valuable — hip abductor weakness is mechanistically linked to poor control of hip adduction during stance, so a weak result on HHD testing supports/explains what you see on the visual closed-chain test.
  • Best practice: use both — HHD gives you an objective, trackable strength number (good for monitoring progress); the closed-chain test shows you if that strength actually translates into functional control during a loaded, dynamic task (which is more mechanically similar to running stance).

Hip flexion angle → goniometer or tabletop test? Yes — goniometer for both active and passive ROM.

  • Active hip flexion ROM (standing or supine, actively lifting the leg) is more functionally relevant than passive ROM for a running-based parameter.
  • Thomas test is useful too, but that's assessing hip flexor tightness/extension deficit, which is a different (though related) construct — worth doing alongside, not instead of, direct flexion ROM measurement.
  • Again, static ROM ≠ dynamic angle during stance/swing phase of running, but restricted active ROM is a sensible flag that dynamic hip flexion during gait may also be altered.

Overall pattern worth naming explicitly: All of these clinical tests measure structural/strength/ROM capacity (what the joint can do), while the paper's parameters measure task-specific dynamic behavior during running (what the joint actually does under load, at speed, repetitively). They correlate, but not perfectly — someone can have full hip ROM on a table yet still run with excessive hip adduction due to fatigue, poor motor control, or compensations elsewhere in the chain.

For a more direct bridge to what this paper actually measured, the video gait analysis (sagittal + posterior view, treadmill running) remains your most valid low-cost option — the clinical tests above are best used as complementary screens to explain why a gait deviation might be happening (weakness, restricted ROM) rather than as substitutes for observing the deviation itself.


Let's try this again. For my single leg jump question, i was implying on force plates

C

Got it — that changes things meaningfully since force plates give you actual kinetic data, not just a qualitative impression.

Single-leg hop/jump test on force plates → Ankle joint reaction force

This is a legitimately strong proxy, closer to the spirit of what the paper measured (kinetic loading) than any of the ROM/goniometer tests.

What force plates can give you directly:

  • Peak vertical ground reaction force (vGRF) on landing — a well-established correlate of impact loading through the lower limb, including the ankle/tibia.
  • Loading rate (how quickly force rises after contact) — arguably more clinically relevant than peak force alone, since high loading rates are strongly associated with tibial stress injury risk in the existing literature (independent of this paper).
  • Time to peak force / contact time — stiffer, less absorptive landings (short contact time, fast rise to peak) suggest higher effective joint loading, including at the ankle.
  • Limb symmetry index — comparing GRF metrics between limbs is one of the more clinically actionable outputs, since side-to-side asymmetry is a recognized injury risk marker in its own right.

Important nuance — GRF vs. ankle JRF specifically Ground reaction force is the external force at the foot-ground interface. Ankle joint reaction force (what the paper actually modeled) is the internal force transmitted through the ankle joint, which is influenced by GRF but also by muscle forces (e.g., triceps surae contraction) acting across the joint — it's typically larger than GRF because muscle forces add to it. Force plates alone can't isolate that internal joint force without inverse dynamics/musculoskeletal modeling (which is exactly the computationally expensive step this paper was trying to avoid).

So:

  • Peak vGRF and loading rate are good, well-validated, accessible proxies you can act on clinically.
  • They're not numerically equivalent to ankle JRF, but they move in the same direction — someone with high impact GRF and fast loading rates is very plausibly also generating high ankle joint reaction forces, especially if they show poor eccentric ankle/calf control on landing (excessive rearfoot strike, minimal ankle dorsiflexion excursion on impact, "stiff" landing strategy).

Practical protocol suggestion

  • Single-leg drop landing or single-leg hop-for-distance with stick landing, on a force plate (or portable force plate system like a Bertec or ForceDecks, if your clinic has one)
  • Capture: peak vGRF, loading rate, contact time, and limb asymmetry index
  • Pair with video (sagittal view) to visually assess ankle dorsiflexion excursion during landing — a stiff-ankle landing strategy (minimal dorsiflexion, rapid loading) is the mechanistic story connecting your force plate numbers to elevated ankle joint loading

That combination — force plate kinetics + visual ankle strategy — gets you meaningfully closer to the construct the paper cared about than any static/manual test would.


Single leg hop for distance is not an option. Is single leg drop landing the best test on a force plate? Also, if muscle forces are important, would a plantar flexion isometric push test (on force plate) be helpful as an adjunct? Whether it be standing, seated or in half kneeling?

C

Good follow-up questions — let's refine both.

Is single-leg drop landing the best force-plate option?

It's a solid choice, but there are a few variants worth considering depending on what you're trying to capture:

  • Single-leg drop landing (stick landing, no rebound) — best for assessing pure landing/absorption strategy. Gives you peak vGRF, loading rate, contact time. Good if your goal is isolating deceleration control.
  • Single-leg drop jump (drop landing → immediate rebound jump) — arguably more relevant to running than a stick-landing drop, because running is a reactive, cyclic loading task, not a single absorb-and-stop event. This lets you calculate:
  • Reactive Strength Index (RSI) = jump height ÷ ground contact time. RSI is a well-established metric of how efficiently someone uses the stretch-shortening cycle (i.e., tendon/muscle stiffness and reactive capacity at the ankle) — which is mechanistically close to what happens at each footstrike during running.
  • This is probably your best single force-plate test if you had to pick one, since it captures both loading (impact) and the reactive/muscular response (push-off), which together approximate the dynamic story behind ankle joint loading better than a static drop landing alone.
  • Repeated hopping in place (single-leg, e.g., 10 seconds continuous) — most closely mimics the cyclic, high-frequency loading pattern of running gait. Gives you average contact time, RSI, and vGRF across multiple cycles rather than one snapshot — arguably the most ecologically valid of the three, but takes more setup/analysis effort.

Bottom line: if simplicity is the goal, single-leg drop landing is fine. If you want something closer to the running-specific loading pattern, a single-leg drop jump with RSI calculation is a meaningful upgrade and still very feasible on a force plate.


Would a plantarflexion isometric push test be a useful adjunct?

Yes — good instinct, and it addresses a real gap: the drop landing/jump tests capture net force absorption, but not how much of the ankle joint force is coming from active muscle contribution (triceps surae) versus passive/structural loading. An isometric plantarflexor test adds that muscle-capacity piece.

A few things to consider:

Positioning matters — it changes which muscle you're testing:

  • Standing (knee extended) — engages both gastrocnemius and soleus. This most closely resembles the ankle position and loading during mid-to-late stance in running, so it's probably the most functionally relevant position for your purposes.
  • Seated or half-kneeling (knee flexed ~90°) — biases toward the soleus, since gastrocnemius is slackened when the knee is flexed. This is clinically useful because the soleus is the primary muscle controlling tibial loading during stance phase of running (much more so than gastrocnemius, which is more active in propulsion/push-off). If you're specifically interested in the muscle contribution to tibial stress fracture risk, a knee-flexed isometric plantarflexion test may actually be the more mechanistically relevant adjunct.

Practical suggestion: test both if time allows — standing (gastroc + soleus, functionally global) and half-kneeling or seated (soleus-isolated). The soleus-specific number is probably the more clinically meaningful one given what you're trying to predict, but the standing test still has value since it reflects the whole plantarflexor unit's contribution during stance.

One limitation to flag: isometric strength (peak force in a static, controlled position) doesn't fully represent dynamic, high-velocity muscle force production during running stance, which happens in a stretch-shortening cycle context, not an isometric one. So:

  • The isometric test tells you strength/force capacity (how much force the muscle can produce).
  • The drop jump/RSI test tells you how well that capacity is actually being utilized reactively during a dynamic, ballistic task.

Putting it together: a runner with strong isometric plantarflexion force but poor RSI on the drop jump suggests a "capacity without utilization" problem — likely motor control, timing, or tendon stiffness issue rather than pure weakness. A runner weak on both suggests a true strength deficit driving poor load absorption. That combination gives you a genuinely useful clinical picture that's more sophisticated than either test alone, and reasonably close (within clinical/practical limits) to the mechanistic story the paper is describing with ankle joint reaction force.


Ok so single leg drop + rsi vs repeated hopping in place

C

Direct comparison, since you're choosing between them:

Single-leg drop jump + RSI

Pros:

  • Quick, single-trial test (or 3 trials averaged) — low time cost, easy to fit into a session
  • Well-established protocol with normative reference values in sports med/S&C literature
  • Clean, interpretable RSI number
  • Standardized drop height controls the loading stimulus, making it more reproducible across visits (good for tracking progress over time)
  • Easier to coach — most patients pick up the "land and immediately jump as high/fast as possible" instruction quickly

Cons:

  • Single event — one good or bad rep can skew the number; less representative of a fatiguing, repetitive task
  • Loading pattern (vertical drop) is a bit different from running (more horizontal progression, shorter ground contact times at higher speeds, different joint angles at contact)
  • Requires a box/platform of consistent height for standardization

Repeated single-leg hopping in place

Pros:

  • Cyclic, rhythmic loading pattern is more mechanically similar to running gait — repeated ground contacts, continuous stretch-shortening cycling
  • Reveals fatigue-related breakdown — you can watch/measure whether contact time lengthens, RSI drops, or asymmetry emerges over the course of 10-15 seconds, which is clinically valuable since stress fractures develop under repetitive, fatiguing load, not a single landing
  • No equipment needed beyond the force plate itself (no box)
  • Can reveal compensatory strategies that only show up after several reps (e.g., someone looks fine on rep 1 but shifts strategy by rep 8)

Cons:

  • More data to process/interpret (multiple contacts vs. one clean number)
  • Harder to standardize test-retest (patients may unconsciously change pacing, height, or rhythm between sessions)
  • More fatiguing/demanding for the patient, which could be a limiting factor with deconditioned or pain-limited patients
  • Slightly more setup/coaching to get a consistent, rhythmic hopping pattern (some patients drift in tempo or height)

My recommendation given your use case (stress fracture risk screening):

Repeated hopping in place is probably the better fit conceptually, because tibial stress fracture risk is fundamentally about cumulative, repetitive loading and how someone's strategy holds up under that — not their capacity for a single well-executed landing. The ability to detect within-test degradation (rising contact time, falling RSI, increasing asymmetry across reps) is arguably the single most clinically meaningful piece of data you could get from a force plate for this specific question, and the drop jump simply can't show you that.

However, if your priority is a clean, standardized, easy-to-track number for monitoring change over time (e.g., in a rehab progression across weeks), the single-leg drop jump is more practical and reproducible.

Best of both worlds if time allows: start with a single-leg drop jump for a clean baseline RSI, then follow with repeated hopping (10-15 reps) to see if that RSI holds up or degrades under repetition. The drop jump gives you the "peak capacity" number; the repeated hopping gives you the "durability under repetitive load" story — which is really the more direct analog to what happens over a 10-week training block in the actual paper.


Let's go with repeated single leg hopping in place. Now summarize the practical takeaways - the tests - for in clinic

C

Practical In-Clinic Testing Protocol

Based on our discussion, here's the consolidated screening battery for the parameters from the paper:


1. Ankle joint reaction force (proxy) Test: Repeated single-leg hopping in place on a force plate (~10–15 seconds continuous) What to capture:

  • Peak vGRF and loading rate per contact
  • Ground contact time (watch for lengthening over reps = fatigue-related breakdown)
  • Reactive Strength Index (RSI) trend across reps (declining RSI = poor reactive/absorptive strategy under repetitive load)
  • Limb symmetry (compare L vs. R) Why: Cyclic, repetitive loading pattern most closely mirrors running gait and reveals how loading strategy holds up under fatigue — the core mechanism behind stress fracture development.

Adjunct: Isometric plantarflexion push test on force plate

  • Half-kneeling or seated (knee flexed ~90°) — soleus-dominant, most relevant to tibial loading during stance
  • Standing (knee extended) — gastroc + soleus, more global
  • Compares muscle force capacity (isometric test) against utilization (hopping RSI) — mismatch between the two tells you whether deficits are strength-based or motor-control/timing-based

2. Ankle Dorsiflexion ROM Test: Weight-bearing lunge test with goniometer (knee-to-wall, measuring tibial angle) Why: Validated, functionally relevant (weight-bearing), quick


3. Hip External Rotation angle (men) Test: Prone or tabletop goniometry, both directions Why: Standard, reliable clinical proxy for rotational hip mobility


4. Hip Adduction Moment (women) Tests (use both together):

  • Hand-held dynamometer — side-lying hip abductor strength (objective, trackable number)
  • Closed-chain visual test — single-leg squat or step-down, watching for hip drop/dynamic valgus Why: HHD tells you strength capacity; closed-chain test tells you whether that strength translates to functional control under load — combination is more clinically informative than either alone

5. Hip Flexion angle (women) Test: Active hip flexion ROM with goniometer (standing or supine) Adjunct: Thomas test (hip flexor tightness/extension deficit) — related but distinct construct


Bridge test (recommended addition): Slow-motion smartphone video (sagittal + posterior views) during treadmill running — captures the actual dynamic gait pattern these static/force-plate tests are only proxying for. Use Kinovea or similar free software for frame-by-frame angle estimates at initial contact/midstance.


Key framing for interpretation: The goniometer/HHD/closed-chain tests measure capacity (what the joint/muscle can do); the force plate hopping test measures behavior under repetitive dynamic load (what it actually does); the video captures real running mechanics. Discrepancies between these layers (e.g., good static ROM but poor dynamic RSI, or good isometric strength but poor closed-chain control) are often the most clinically useful findings — they point toward motor control or fatigue-related deficits rather than pure structural limitations.


Is measuring Hip External Rotation angle the best test based on what the authors found?

C

Good pushback — let's look closer at what the paper actually shows, because the answer is more nuanced than my earlier summary implied.

What the data actually says about hip external rotation angle (men):

  • In the combined kinematic+kinetic model, hip external rotation angle appeared in 45% of the 500 models — the second most frequent parameter behind ankle JRF (99%). That sounds decent.
  • But in the kinematic-only model (Figure 3B), hip external rotation angle dropped to just 14% — tied for the lowest of the top three, behind ankle dorsiflexion angle (57%) and hip adduction angle (14%).
  • The authors explicitly note: "only the ankle dorsiflexion angle appeared in more than 50% of the models, indicating the absence of parameters with strong predictive power" for the men's kinematic-only model.
  • Consistent with this, the LR classification model using ankle dorsiflexion + hip external rotation (kinematic-only, men) performed poorly — AUC only 0.53–0.71, with most confidence intervals crossing 0.50 (meaning not reliably better than chance).

So — is a goniometer for hip external rotation the best test?

Not really, based on this paper's own findings. Here's the honest read:

  • Hip external rotation angle only earned its "top predictor" status when paired with kinetic data (ankle JRF). On its own, as a pure kinematic/ROM measure, it was a weak signal for men — appearing in only 14% of models.
  • Ankle dorsiflexion angle is actually the better-supported kinematic test for men (57% selection frequency) — this is the one I should have emphasized more, not hip rotation.
  • The paper's real message for men is: kinematic parameters alone (including hip rotation) just aren't that predictive — you need the kinetic data (ankle JRF) to get meaningful discrimination. That's exactly why the men's kinematic-only AUCs were so weak.

Revised practical takeaway:

For men specifically, based on the paper's actual evidence strength:

  1. Prioritize the ankle joint reaction force proxy (repeated hopping on force plate) — this is where nearly all the men's predictive power lives (99% selection frequency)
  2. Ankle dorsiflexion ROM (goniometer/weight-bearing lunge) is the better-supported kinematic backup test, not hip rotation
  3. Hip external rotation goniometry — reasonable to include as a minor/adjunct data point, but the paper's own numbers suggest it shouldn't be weighted heavily as a standalone screening test for men; its predictive value was tied to being combined with kinetic (force) data, not on its own

Good catch — this is a case where the "top 2 predictors" framing in the abstract can be a bit misleading if you don't look at how those numbers changed between the combined and kinematic-only analyses.


But would hip internal ROM be useful as well?

C

Fair question — but here the evidence trail actually argues against adding it, at least based on this paper specifically.

What the paper says about hip internal rotation angle:

Hip internal rotation angle was one of the 30 parameters they tested (parameter #14 in Table 2), included in both the full and kinematic-only analyses. It's notably absent from every "top 3" list in the paper — not mentioned as a top predictor for men or women, in either the combined kinematic+kinetic analysis or the kinematic-only analysis. The authors explicitly state that none of the remaining 27 parameters (of the 30 total) appeared in more than 25% of the 500 models for the full analysis — and hip internal rotation was one of those 27. So based on this specific dataset, it showed essentially no meaningful predictive signal for tibial SF risk.

So strictly by the paper's evidence: no, it wouldn't be a priority.

But there's a reasonable clinical argument for measuring it anyway — separate from this paper's findings:

  • Hip internal and external rotation are typically assessed as a pair clinically (e.g., prone or seated goniometry), since they reflect the same joint's rotational capacity and are usually measured in the same test position with minimal added time/cost.
  • Restricted internal rotation ROM can drive compensatory external rotation during gait — if someone lacks IR mobility, they may functionally substitute with excess ER during stance, which is the parameter the paper flagged as predictive in men. In that sense, IR ROM could serve as an explanatory variable for why someone shows the ER pattern, even if IR itself isn't independently predictive of fracture risk.
  • It's essentially free to add once you're already goniometrically testing ER — no real added burden.

Bottom line:

  • Don't treat hip internal rotation as an evidence-based predictor from this paper — it wasn't identified as informative in the actual results.
  • Reasonable to still collect it as context, since it's assessed in the same test position as external rotation and can help you interpret why a patient shows excessive external rotation during running (a capacity/compensation story), rather than treating it as a standalone risk marker in its own right.

This is a good example of separating "what the paper's data supports" from "what's clinically sensible practice" — they're not always the same thing, and it's worth being explicit about which one you're drawing on.


So let's summarize the take home tests again, including what to look for/why its important

C

Final In-Clinic Screening Protocol

Corrected to reflect actual strength of evidence from the paper.


1. Ankle Joint Reaction Force (proxy) — strongest predictor, both sexes

Test: Repeated single-leg hopping in place on a force plate (10–15 sec continuous)

What to look for:

  • Peak vGRF and loading rate per contact
  • Ground contact time — watch for lengthening across reps (sign of fatigue-driven breakdown in absorption strategy)
  • RSI (reactive strength index) trend — a declining RSI over the trial suggests the patient can't sustain efficient load absorption under repetitive stress
  • L vs. R asymmetry

Why it matters: This was the single most predictive variable in the entire paper (selected in 76–99% of models across both sexes). It's also the parameter that best mirrors the actual mechanism of a stress fracture — repetitive submaximal loading that eventually exceeds bone's capacity to repair itself.

Adjunct — isometric plantarflexion push test (half-kneeling/seated for soleus-dominant, or standing for gastroc+soleus) Why: Tells you whether poor hopping performance reflects a strength deficit (weak on both isometric and hopping) or a motor-control/timing problem (strong isometric, poor RSI) — changes your treatment approach.


2. Ankle Dorsiflexion ROM — best kinematic (non-force) test, both sexes

Test: Weight-bearing lunge test with goniometer (knee-to-wall)

What to look for: Reduced available dorsiflexion ROM (compare to normative ~35–40° or side-to-side asymmetry)

Why it matters: This was the most consistently useful kinematic predictor in the paper for both men (57% of kinematic-only models) and women (71%). Restricted dorsiflexion has separately been linked in the literature to altered hip/pelvis mechanics and increased injury risk. Unlike hip rotation, this one holds up as a standalone signal even without kinetic data.


3. Hip Adduction Moment (proxy) — top predictor for women specifically

Tests (pair together):

  • Hand-held dynamometer — side-lying hip abductor strength
  • Closed-chain visual test — single-leg squat or step-down, watching for hip drop/dynamic valgus

What to look for: Weak abductor strength on HHD and/or visible hip drop/valgus collapse during the functional test

Why it matters: Hip adduction moment was selected in 64% of women's models and, paired with ankle JRF, produced the paper's strongest classification performance for women (AUC up to 0.93). This is a well-supported, sex-specific risk marker — prioritize it in female patients.


4. Hip Flexion ROM — top kinematic predictor for women

Test: Active hip flexion ROM with goniometer (standing or supine)

What to look for: Reduced active flexion ROM, or asymmetry between limbs

Why it matters: This was the single strongest kinematic parameter for women (79% selection frequency) — even more consistently predictive than ankle dorsiflexion in women specifically.


5. Hip Rotation ROM (internal + external) — secondary/contextual only, don't overweight

Test: Prone or tabletop goniometry, both directions

What to look for: Restricted internal rotation with compensatory excess external rotation pattern

Why it matters (with caveat): External rotation angle only showed strong predictive value in men when combined with ankle JRF data; alone, it was weak (only 14% selection frequency, AUC dropped to 0.53–0.71 — near chance). Don't treat this as a primary screening test. Still worth measuring as context — poor internal rotation ROM may explain why a patient compensates with excess external rotation during running — but it shouldn't drive clinical decision-making on its own.


Bridge test: Running gait video

Test: Slow-motion smartphone video (sagittal + posterior views) during treadmill running, analyzed with Kinovea or similar

Why it matters: All the above are static/isolated tests serving as proxies. Video lets you confirm whether deficits identified on the table (e.g., limited hip flexion ROM, weak abductors) actually show up as altered mechanics during the dynamic task that matters — running.


Interpretation framework

Layer What it tells you
Goniometer/HHD (static) Structural capacity — what the joint/muscle can do
Force plate hopping (dynamic) Behavior under repetitive load — what actually happens with fatigue
Gait video (functional) Real running mechanics — whether deficits translate to the task itself

A mismatch between layers (e.g., normal ROM but poor RSI decay, or good isometric strength but visible hip drop on functional testing) is often more clinically useful than any single number — it points you toward motor control, fatigue resistance, or movement-pattern issues rather than pure structural limitation.


Please re-write and include any important requisite baseline and/or normative data

C

In-Clinic Screening Protocol — with Baseline/Normative Reference Data

Important framing before the tests: this paper did not establish validated clinical cutoffs — it identified which parameters matter, not what values indicate danger. So normative data below comes from the broader clinical/biomechanics literature, not from this paper itself. Where norms are weak or population-specific, that's flagged. Side-to-side asymmetry and within-patient baseline tracking are generally more clinically reliable than comparing to population norms.


1. Ankle Joint Reaction Force (proxy) — strongest predictor, both sexes

Test: Repeated single-leg hopping in place on a force plate (10–15 sec continuous)

What to look for / reference data:

  • Contact time: should remain roughly stable across the trial; a progressive increase of >10–15% from early to late reps suggests fatigue-driven breakdown in absorptive strategy
  • RSI: no established norm for this specific population (military recruits/runners); trained athletes typically show RSI values in the ~1.5–2.5+ range, deconditioned individuals lower — treat this as a within-patient baseline metric (track over time / compare limbs) rather than benchmarking against generic athletic norms
  • Limb symmetry index (LSI): <85–90% on any metric (peak force, RSI, contact time) is a commonly used red-flag threshold in return-to-sport literature
  • Establish a baseline at intake — this is your most valuable reference point, since normative running/hopping kinetic data is sparse and population-dependent

Why it matters: Most predictive variable in the entire paper (76–99% model selection frequency, both sexes); best mirrors the actual mechanism of repetitive submaximal overload central to stress fracture development.

Adjunct — isometric plantarflexion push test (half-kneeling/seated = soleus-dominant; standing = gastroc+soleus)

  • No strong universal normative value; compare limb-to-limb (LSI <85–90% flagged) and track over time
  • Why: Differentiates strength deficit (weak on both isometric and hopping) from motor-control/timing deficit (strong isometric, poor RSI)

2. Ankle Dorsiflexion ROM — best kinematic (non-force) test, both sexes

Test: Weight-bearing lunge test (knee-to-wall), goniometer or inclinometer

Reference values:

  • Normal tibial angle: ~35–45°
  • Knee-to-wall distance: ~9–12 cm typically considered adequate
  • Clinically meaningful restriction/asymmetry: side-to-side difference >4–5° or >1.5 cm is commonly used as a flag in the sports medicine literature

Why it matters: Most consistently useful kinematic predictor in the paper — 57% selection frequency in men, 71% in women.


3. Hip Adduction Moment (proxy) — top predictor for women specifically

Tests (pair together):

  • Hand-held dynamometer — side-lying hip abductor strength
  • Closed-chain visual test — single-leg squat or step-down

Reference values:

  • HHD: normalize to body weight (force ÷ body weight); absolute norms vary widely by device/positioning, so limb symmetry <85–90% is the more actionable cutoff than any absolute number
  • Single-leg squat/step-down: use a simple qualitative grading scale (e.g., good / fair / poor hip control), looking for pelvic drop, femoral adduction, or knee valgus — no universal numeric norm, but visible frontal-plane collapse is a well-established qualitative red flag in the literature

Why it matters: Selected in 64% of women's models; combined with ankle JRF, produced the paper's best classification performance for women (AUC up to 0.93). Highest-priority, sex-specific test for female patients.


4. Hip Flexion ROM — top kinematic predictor for women

Test: Active hip flexion ROM, goniometer (standing or supine)

Reference values:

  • Normal active hip flexion: roughly 100–120° (varies by source/method; passive ROM norms run slightly higher, ~120–135°)
  • Side-to-side asymmetry >10° is a reasonable clinical flag

Why it matters: Single strongest kinematic parameter for women in the paper (79% selection frequency) — more consistently predictive than ankle dorsiflexion in women specifically.


5. Hip Rotation ROM (internal + external) — secondary/contextual only

Test: Prone or tabletop goniometry, hip neutral, knee flexed 90°

Reference values:

  • Normal internal rotation: ~30–40°
  • Normal external rotation: ~40–60°
  • Asymmetry >10°, or a notable IR/ER imbalance (e.g., low IR with compensatory high ER), is the more clinically meaningful finding than either number in isolation

Why it matters (caveat): External rotation angle only showed strong predictive value in men when combined with kinetic data; alone, it was weak (14% selection frequency, AUC only 0.53–0.71, often not better than chance). Don't treat as a primary screening test — use it as context for interpreting compensatory patterns, not as a standalone risk flag.


Bridge test: Running gait video

Test: Slow-motion smartphone video (sagittal + posterior views), treadmill running, analyzed via Kinovea or similar

No standard norms exist for angle-at-contact values from smartphone video — this test's value is qualitative and comparative: does what you found on the table (restricted ROM, weak hip abductors, poor RSI decay) actually show up as a visible deviation during running? Establish a baseline video at intake to track change over a rehab course.


Summary Table

Test Key metric Reference/flag Priority
Single-leg repeated hopping (force plate) Contact time drift, RSI trend, LSI LSI <85–90%; contact time ↑10–15% over trial Highest — both sexes
Isometric plantarflexion push Peak force, LSI LSI <85–90% Adjunct to above
Weight-bearing lunge (ankle DF) Tibial angle / knee-to-wall ~35–45° / ~9–12cm; asymmetry >4–5° or >1.5cm High — both sexes
HHD hip abductor + single-leg squat Force (BW-normalized), visual quality LSI <85–90%; visible pelvic drop/valgus High — women
Active hip flexion ROM Degrees ~100–120°; asymmetry >10° High — women
Hip IR/ER goniometry Degrees IR ~30–40°, ER ~40–60°; asymmetry >10° Low/contextual only
Running gait video Qualitative Baseline comparison over time Confirmatory, all patients

Key caution: Because this paper didn't validate specific clinical thresholds, and most of the reference values above come from general orthopedic/sports medicine literature rather than a stress-fracture-specific validated database, within-patient tracking over time and side-to-side symmetry are your most defensible clinical anchors — more so than comparing any single patient to a population "normal."


End of transcript.

0 Comments



Leave a Reply.

    Picture

    Jeff Cubos

    MSc, DC, FRCCSS(C), CSCS

I created this blog to share my thoughts with others. It is not intended to be used for medical diagnosis, medical treatment or to replace evaluation by a health practitioner. If you have an individual medical problem, you should seek medical advice from a professional in your community. Any of the images I do use in this blog I claim no ownership of.
  • Blog
  • Reviews
    • CE Reviews
    • Research Reviews
    • Book Notes
  • About