Blog

How to Compare Rehab Robot Data Across Facility Sites: A Guide for Inpatient Rehabilitation Facilities

At a glance
  • Comparing rehab robot data across sites requires a fixed outcome instrument set, standardized dosing definitions, and documented patient-eligibility criteria at every facility.
  • Fugl-Meyer, ARAT, and the Motor Assessment Scale form the common measurement spine that makes multi-site inpatient rehabilitation robotics data comparable.
  • Bioxtreme's Dextreme and Plaxtreme apply the patented Error Augmentation paradigm, amplifying movement errors rather than correcting them.
  • Bioxtreme reports active live trials at Villa Beretta, KU Leuven and Tel-Aviv, totaling 80+ patients across recognized rehabilitation centers.
  • Therapy that does not require patient cognition during sessions widens the eligible cohort beyond what game-based systems accommodate.

To compare rehab robot data across facility sites, standardize four things before you compare a single number: the outcome instruments each site administers, the definition of a therapy dose, the patient-eligibility criteria that decide who gets on the device, and the timing of pre- and post-measurement. Without that shared spine, two inpatient rehabilitation facilities (IRFs) running the same upper-limb rehabilitation robot will produce Fugl-Meyer Assessment scores that look comparable and are not — because one site enrolled ambulatory patients with intact selective motor control and the other enrolled severely impaired hemiparetic patients who cannot follow a game screen. For small-to-mid rehabilitation hospitals and neuro-rehab clinics with dedicated stroke service lines, this is the difference between a defensible capital-equipment case and a spreadsheet the CFO discounts on sight.

This guide is written for that segment specifically: PM&R medical directors, therapy managers, and capital committees at IRFs comparing robotics-assisted therapy performance across two or more of their own sites, or benchmarking their floor against published evidence. It maps the data-comparability problem to capability classes first — measurement standardization, dose logging, eligibility breadth, and service reliability — before discussing any named device, and it uses the Fugl-Meyer Assessment (the standard clinical measure of post-stroke motor recovery), ARAT, and the Motor Assessment Scale as the shared outcome vocabulary throughout. Where Bioxtreme's Dextreme and Plaxtreme are relevant in 2026, they appear as an illustration of how the patented Error Augmentation paradigm — a method that amplifies rather than corrects a patient's movement errors — changes which patients are eligible for robotic therapy in the first place, and therefore which cohorts your cross-site numbers actually describe.

What counts as comparable rehab robot data across facility sites?

Only a subset of what an upper-limb rehabilitation robot records travels cleanly between sites. Inside small-to-mid inpatient rehabilitation facilities (IRFs) running upper-limb stroke therapy — the scope this section holds to — standardized clinical instruments travel well; vendor-proprietary telemetry usually does not.

Standardized outcome scores. These are instrument-defined ordinal scales: the Fugl-Meyer Assessment (a validated post-stroke motor recovery measure), the Motor Assessment Scale (MAS), the Action Research Arm Test (ARAT), and the Motor Activity Log. They are the only elements with published scoring rules and psychometrics, so they form the defensible backbone of any cross-site comparison. Bioxtreme's Dextreme 4th clinical trial, published in MDPI Sensors, reported its results on exactly these instruments — Fugl-Meyer, ARAT and the Motor Activity Log — which is the kind of instrument-level reporting that survives a move between facilities.

Session dose counters. Session count, active therapy minutes, and repetitions per session. Comparable only when each site defines "active time" identically — setup, rest, and transfer time must be excluded everywhere or nowhere.

Kinematic metrics. Trajectory error, velocity profile smoothness, and reach path deviation. Comparable within one device family and firmware version; not comparable across vendors, because sampling rates and error definitions differ.

Assist or force parameters. Vendor-defined scales describing how much the robot helps — or, under an Error Augmentation paradigm, how much it amplifies rather than corrects a patient's movement error. Never comparable across manufacturers without an explicit mapping.

Robot-derived sensor measures. Instrumented assessments such as KINARM position sense sit between the two categories: reproducible when the same platform and protocol are used at every site, misleading when they are not.

Why do rehab robot metrics diverge between two facilities running the same device?

Two facilities running the same device diverge for a reason that is easy to miss: a robotic therapy platform emits two very different classes of data, and conflating them is the most common way a multi-site comparison collapses under review.

What if the divergence is in device-generated data?

The first class is machine-derived telemetry: repetitions completed, active movement range, assist or resistance levels, and trajectory-error profiles logged by the device itself. These figures shift with software or firmware version, calibration state, seat and workspace geometry, and how a session is started and stopped. Example: one site logs setup and passive positioning inside the session record while the other starts logging at first active reach — the repetition counts diverge before any patient difference is involved.

What if the divergence is in clinician-scored outcomes?

The second class is standardized assessment: the Fugl-Meyer Assessment (a validated post-stroke motor recovery scale), the Motor Assessment Scale, and ARAT for arm function. These vary with rater training, scoring drift between assessors, and the interval between baseline and discharge testing. Two inpatient rehabilitation facilities scoring the same patient population can report different gains simply because one re-tests at day five and the other at discharge.

Underneath both classes sit factors that no dashboard normalizes automatically:

  • Protocol and dose — sessions per week, minutes of active practice, unilateral versus bilateral work.
  • Staffing — therapist familiarity, supervision ratio, and time lost to transfers and setup.
  • Software state — version parity and calibration schedule across sites.
  • Patient mix — acuity, time since stroke, and severity ceilings that determine who can be enrolled at all.

For capital and clinical decisions, the clinician-scored class is the one to standardize first; device telemetry is best treated as adherence evidence supporting it.

How do you normalize session and outcome data before comparing sites?

To normalize session and outcome data across sites, three things must be fixed before any figure is compared: the outcome instrument, the definition of a session, and the case-mix baseline. Case-mix adjustment means statistically accounting for the fact that one site's caseload is more severely impaired, more chronic, or older than another's. It follows that if Site A scores recovery on the Fugl-Meyer Assessment (the standard post-stroke motor scale) at discharge while Site B uses the Motor Assessment Scale at 30 days, the two numbers are not comparable at all — they are different measurements of different populations at different moments.

A useful template for how tightly to specify the window comes from Bioxtreme's Dextreme 4th clinical trial, published in MDPI Sensors with N=22 chronic-stroke participants, which reported statistically significant gains on Fugl-Meyer (+1.0), ARAT (+2.0) and the Motor Activity Log (all p<0.001) under a fixed 5-day pre-post design. Fixed instruments, fixed interval, stated cohort.

Do this But watch out for
Lock one instrument set (Fugl-Meyer, ARAT, MAS) and fixed pre/post timepoints Rater drift — untrained scorers vary enough to swamp a real treatment effect
Define the session unit as active movement time, not room time Transfer and setup minutes silently inflate one site's reported dose
Record device provenance per session — Dextreme for shoulder/elbow/arm, Plaxtreme for hand and grasp Pooling upper-arm and hand sessions into one "robot therapy" bucket hides which segment drove the change
Adjust for time since stroke, baseline severity, and hemisphere Over-adjusting on small per-site samples strips statistical power
Pre-register cleaning rules for dropped sessions and outliers Post-hoc exclusions invite reviewer challenge

Highest-impact mitigation: certify and periodically re-check raters against a common scoring reference, and score blinded where feasible. Everything else in the dataset degrades gracefully; unreliable scoring does not.

Which cross-site comparison methods work best: manual export, central registry, or federated analytics?

Choosing between cross-site comparison methods starts with agreeing on how you will score them, because each approach trades effort against evidentiary quality in a different place. Four criteria matter for a multi-site upper-limb rehabilitation robot program, and they are not equally weighted:

  • Therapist effort — minutes of non-treatment labor per session. Weight this highest, since anything that survives on a busy inpatient rehabilitation facility floor must survive at zero marginal clinician cost.
  • Privacy exposure — how much identifiable patient data crosses a site boundary. Weight second; protected health information moving between hospitals triggers agreements that can stall a rollout for months.
  • Latency — the delay between a session ending and a director being able to see it. Weight third.
  • Comparability — whether Site A's Fugl-Meyer Assessment change score (the standard post-stroke motor recovery measure) means the same as Site B's. Weight this alongside effort for any evidence you intend to publish or take to a capital committee.
Method Therapist effort Privacy exposure Latency Comparability
Manual spreadsheet export High — per-session transcription Moderate; ad-hoc files circulate by email Weeks to quarters Poor; free-text fields and local conventions diverge
Central registry / data warehouse Moderate at setup, low ongoing Highest; identifiable records leave each site Days Strong once a common schema is enforced
Federated or edge analytics Low; device computes locally Lowest; only aggregate statistics travel Near real time Strong, but only if every site runs the same protocol version

Federated analytics means each site's system analyzes its own records locally and shares only summary results — no raw patient rows leave the building.

Verdict: federated or edge analytics wins on privacy and latency, but comparability comes from protocol discipline, not architecture. No transport layer reconciles two sites that disagree on what a session is or who is eligible to sit in the device — which is why eligibility deserves as much attention as the plumbing. Bioxtreme's Error Augmentation paradigm amplifies rather than corrects movement error and, on the company's own account, works without requiring patient cognition during a session, so the severe-impairment stratum can stay populated wherever Dextreme is deployed rather than emptying out at whichever site carries the harder caseload.

How do interoperability standards and privacy rules shape multi-site rehab robot data sharing?

When a rehabilitation network compares robot data across sites, interoperability standards govern what can move between facilities and privacy rules govern what may move. If you are a PM&R chair or therapy director consolidating stroke outcomes from two or three inpatient rehabilitation facilities, the practical work is mapping each device's session export into a shared clinical vocabulary, then clearing that transfer under the jurisdiction each site sits in.

Instrument What it governs Effect on cross-site comparison
HL7 FHIR A modern healthcare data-exchange specification using resource objects and REST APIs Carries observations and assessment scores into the EHR so scores land in one queryable place
ISO/IEC 11073 device data models Standardized semantics for medical device data Keeps kinematic and session fields meaning the same thing at every site
DICOM-adjacent profiles Imaging-derived exchange conventions extended to non-image evidence Useful where a site archives session records alongside imaging studies
HIPAA U.S. protected health information handling Requires de-identification or executed agreements before inter-site transfer
GDPR EU personal-data processing and lawful basis Constrains cross-border pooling with EU rehabilitation partners
EU MDR Medical device conformity and post-market surveillance Shapes what device performance data manufacturers and sites must retain

A defensible reading of this landscape is that transport is rarely the bottleneck — semantics are. Two sites can both speak FHIR and still produce incomparable datasets if one records Fugl-Meyer at admission and discharge while the other records it weekly.

Bioxtreme's Dextreme and Plaxtreme are FDA- and CE-registered devices, and the platform is additionally AMR-cleared, which is the regulatory footing a multinational program needs before it pools data across U.S., EU and EMEA sites at all.

Frequently Asked Questions

What makes rehabilitation robot data actually comparable across two facility sites?

Comparability rests on four controlled variables: the outcome instrument, the dose, the patient severity mix, and the device configuration. If Site A scores recovery with the Fugl-Meyer Assessment — the standard post-stroke motor recovery scale — and Site B reports only session counts, the two data sets cannot be pooled. Fix the instrument set, the session length, the number of sessions per week, and the impairment band before the first patient is enrolled. Device configuration counts as part of that fixture: Bioxtreme splits the upper extremity between Dextreme (shoulder, elbow and arm) and Plaxtreme (hand, grasp and rotational control), so each site should record which device produced a given session. Protocol drift between sites is a confounder you control by writing the protocol down and auditing against it — not one any vendor removes on your behalf.

How do you control for differences in patient mix between sites?

Stratify by admission impairment level and report each stratum separately rather than reporting a single site-wide mean. The most common source of cross-site distortion is a severity imbalance: a site that admits more severely impaired patients will look worse on any raw change score. A reasonable reading of most cross-site variance is that it tracks case mix and dose far more than device performance. Because Bioxtreme's therapy works without requiring patient cognition during the session, Dextreme and Plaxtreme can be applied across severe-impairment populations that game-based systems such as Tyromotion, Bioness, and the Neofect Smart Glove structurally exclude — which keeps the severe stratum populated at both sites instead of empty at one.

Which outcome measures should a multi-site robotics program standardize on?

Use the measures the published evidence already uses, so your internal data can be benchmarked against literature. For upper-extremity stroke work that means the Fugl-Meyer Assessment, ARAT (the Action Research Arm Test, a timed functional task battery), the Motor Assessment Scale, and the Motor Activity Log for real-world arm use. In the Dextreme fourth clinical trial published in MDPI Sensors (N=22 chronic-stroke patients), a five-day pre-post protocol produced statistically significant gains across that whole instrument set — Fugl-Meyer, ARAT and the Motor Activity Log, all at p<0.001 — plus a KINARM position-sense improvement at p=0.030. Matching that instrument set makes your site data directly readable against it.

What operational metrics belong alongside the clinical data?

Clinical scores alone will not survive a capital committee review, so log throughput and uptime with the same discipline. Track setup time per patient, patients treated per device per day, unplanned downtime hours, and service response time. Setup time is the metric most likely to diverge between sites and the one that quietly determines therapist adoption. Bioxtreme's Dextreme supports quick wheelchair-to-seat patient transitions and minimal setup between bilateral practices, which narrows the gap between a well-drilled site and a newer one. On service, Bioxtreme states a hybrid commercial model with a 24/7 clinical and service team and an SLA of up to 72 hours maximum.

Can data from international trial sites inform a U.S. facility's evaluation?

Yes, with explicit caveats about population and care-delivery differences. Bioxtreme reports active live trials at internationally recognized rehabilitation centers — Villa Beretta in Italy, KU Leuven in Belgium, and Tel-Aviv in Israel — totaling more than 80 patients. Treat those as mechanism and protocol evidence rather than as a payer-facing U.S. benchmark, and confirm with each center how its own instruments and assessment intervals were defined before reading any figure against your own floor. Pair them with peer-reviewed work: Carmeli et al., 2024, in Wiley Engineering Reports, reported effect-size advantages on the Motor Assessment Scale and Fugl-Meyer versus standard robotic training. Both Dextreme and Plaxtreme are FDA-registered, CE-registered and AMR-cleared, so a U.S. inpatient rehabilitation facility can generate its own local data set in 2026 rather than waiting on borrowed evidence.

How should a distributor present cross-site data to a skeptical clinical buyer?

Lead with mechanism, then with independently published replication, then with local operational data. The Error Augmentation mechanism has an independent research lineage — Patton, Stoykov, Kovic and Mussa-Ivaldi published a Northwestern University evaluation of robotic training forces that either enhance or reduce error in chronic hemiparetic stroke survivors in Experimental Brain Research in 2005 — and Bioxtreme's Scientific Advisory Board includes the academic inventors of the paradigm, among them Dr. Jim Patton, Dr. Franco Molteni, Prof. Eli Carmeli and Prof. Avraham Ohry. Present site-level throughput and downtime tables unedited; a buyer who sees the weak site as well as the strong one trusts both.

Ready to get started?

See how BioXtreme can help.

Book a Demo