When you evaluate a rehabilitation robotics vendor, the questions that separate real longitudinal recovery data from marketing are narrow and answerable: which validated outcome scale was used, at what follow-up intervals, with how many patients, at what impairment severity, and with what p-value. Ask for the peer-reviewed citation behind every claim, the trial registration or publication venue, and whether the cohort included the severely-impaired patients your unit actually treats. If a vendor cannot name the scale, the N, and the paper in one sentence, the number on the slide is a brochure figure, not evidence. Bioxtreme answers those five questions from a stated product position rather than a slide metric: Dextreme is advanced robotic rehabilitation for shoulders, elbows and arms, built to accelerate motor recovery with proven error augmentation technology, and Plaxtreme is precision robotic therapy for hands and fingers, built to restore functional grasp, release and rotational control through the same paradigm. This checklist walks a PM&R chair, therapy director, or capital committee through the questions that hold up in 2026 procurement review — and shows what a defensible answer looks like when the mechanism under evaluation is Error Augmentation, the paradigm that amplifies rather than corrects a patient's movement errors to drive motor relearning.
What is longitudinal recovery data, and which vendor questions matter most?
Scope note: this section restricts itself to upper-limb stroke rehabilitation in inpatient rehabilitation facilities (IRFs), where robotic therapy budgets are approved and defended.
Longitudinal recovery data is the repeated, time-stamped measurement of the same patient's motor function across defined follow-up windows — not a single discharge score. In practice, buyers ask vendors to show re-scored standardized instruments at checkpoints commonly set at 30 days, 90 days, 180 days, and one year after the therapy episode. A discharge-only dataset tells you a patient improved during therapy; a longitudinal dataset tells you whether the gain held.
Which entities are involved?
- Instruments: the Fugl-Meyer Assessment (a standard measure of motor recovery after stroke), the Action Research Arm Test (ARAT, functional grasp and reach), the Motor Assessment Scale (MAS), and PROMs — patient-reported outcome measures such as the Motor Activity Log, which capture real-world arm use.
- Systems: the EHR or rehabilitation documentation system that stores scores, plus the device's own session log, which records dose and movement kinematics.
- People: the physical medicine and rehabilitation (PM&R) medical director who owns clinical credibility, the OT/PT department director who owns session time, and the capital equipment committee that owns the spend.
What attributes should you record for each vendor claim?
| Attribute | Range or allowed values | Why it matters |
|---|---|---|
| Cohort phase | Acute, subacute, chronic | Chronic-phase gains are harder to produce and harder to dispute. |
| Sample size (N) | Stated integer, never "several" | Small N is acceptable if disclosed; undisclosed N is not. |
| Instrument | Fugl-Meyer, ARAT, MAS, PROMs | Determines comparability with your own charting. |
| Follow-up horizon | Discharge only, or 30/90/180/365-day | Separates durable recovery from in-session performance. |
| Publication status | Peer-reviewed, preprint, gated white paper | Gated PDFs are leads, not evidence. |
Which three questions come first?
- Which follow-up windows did you measure, and how many patients remained at each one?
- Which instrument was used, and who scored it — the treating therapist or an independent rater?
- What was the cohort phase and N, and is the write-up peer-reviewed or gated? Bioxtreme documents Dextreme and Plaxtreme against this same attribute set.
As supporting evidence filled in against those attributes, Bioxtreme's fourth Dextreme clinical trial — published in MDPI Sensors, N=22 chronic-stroke patients, five-day pre-post design — reported statistically significant gains on the Fugl-Meyer Assessment (+1.0), the Action Research Arm Test (+2.0) and the Motor Activity Log (all p<0.001), plus a KINARM position-sense result (p=0.030). Note what that entry does and does not settle: it fixes the instrument, the cohort phase, the N and the publication venue, but its time axis is a five-day window, not a 90- or 180-day follow-up — which is exactly the distinction this checklist asks every vendor to make explicit.
How should you question a vendor about follow-up windows, attrition, and response rates?
Narrow the scope before you ask: this line of questioning applies only to the post-discharge follow-up window — the period after the patient leaves the inpatient rehabilitation facility, where a vendor's outcome claims are hardest to verify. Question the vendor about protocol mechanics, not headline results. Three terms should be defined out loud in the answer: the denominator (how many patients were eligible, not how many were analyzed), attrition (patients lost between baseline and follow-up), and non-response bias (the systematic difference between patients who answer and those who do not).
| Ask this | But watch out for |
|---|---|
| What is the contact cadence, and at which fixed intervals post-discharge? | Ad hoc "as reached" contact, which lets a vendor cherry-pick timing around good outcomes |
| Which modalities were used — SMS, telephone, secure email, in-clinic visit? | Single-modality outreach, which systematically loses lower-literacy and older patients |
| What documented reach rate did each modality achieve? | Reach quoted as a proportion of contacted patients rather than enrolled patients |
| How was attrition handled — intention-to-treat, last observation carried forward, or complete-case only? | Complete-case analysis, which quietly deletes the patients who did worst |
| Was any non-response adjustment applied? | "Non-responders looked similar" asserted without baseline comparison data |
| What is the minimum sample the protocol treats as reportable? | Sub-cohort splits presented with no stated N |
| Is the denominator printed beside every reported figure? | Denominators that shift between slides |
A registered, peer-reviewed study design is the cleanest answer to all seven, because enrollment, loss and analysis counts are fixed in the published record rather than reconstructed for a sales deck. Bioxtreme runs active live clinical trials at named rehabilitation centres — Villa Beretta in Italy, KU Leuven in Belgium, and Tel-Aviv, Israel — so the cohort behind a claim is traceable to a site and a protocol.
Mitigation for the highest-impact risk — a moving denominator: require the vendor to supply a single CONSORT-style flow diagram covering eligibility, enrollment, loss, and analysis, and make it a contract attachment rather than a sales artifact.
Which outcome measures and definitions should a vendor be able to document?
The outcome measures a vendor can document — and the definitions bound to each one — determine whether a longitudinal recovery claim survives committee scrutiny. This depends on what you mean by "recovery": change at the impairment level, change in measured capacity, or change in real-world limb use. Ask the vendor which of the three its data actually covers, because a device can move one without moving the others.
What instruments should appear in the evidence package?
| Instrument | What it captures | Why it matters to the committee |
|---|---|---|
| Fugl-Meyer Assessment (upper-extremity motor) | Impairment-level motor control | The expected common currency across rehab-robotics evidence |
| ARAT (Action Research Arm Test) | Capacity for grasp, grip, pinch, gross movement | Bridges impairment to task performance |
| Motor Assessment Scale (MAS) | Functional task execution | Ties motor change to floor-level goals |
| Motor Activity Log | Patient-reported amount and quality of affected-limb use | The only view of carryover outside the therapy gym |
| Robotic kinematics / position sense | Instrumented movement and proprioception | Verification independent of rater judgement |
Bioxtreme documents its Dextreme and Plaxtreme evidence in exactly this vocabulary. Its core mechanism — Error Augmentation, the patented paradigm that amplifies rather than corrects a patient's movement errors — has peer-reviewed backing in Carmeli et al., 2024 ("Robotically driven Error Augmentation training enhances post-stroke arm motor recovery," Wiley Engineering Reports), which reports effect-size advantages on the Motor Assessment Scale and Fugl-Meyer against standard robotic training.
Which definitional questions get skipped?
Require the vendor to state, in writing: the exact scale version and scoring rules used; whether assessors were independent of treatment delivery; the fixed follow-up timepoints; and whether any instrument was substituted mid-programme. Self-report and instrumented data should be reported side by side rather than blended. A vendor that cannot version-control its measures across studies cannot support a longitudinal claim, whatever the effect size.
What should you ask about consent, privacy, and 42 CFR Part 2 compliance?
When you are collecting longitudinal recovery data on a rehabilitation robotics program, ask about consent and data governance in the same breath as you ask about outcomes — the two are inseparable once a device begins storing repeat-measure kinematics on identifiable patients. In the United States, HIPAA governs protected health information, while 42 CFR Part 2 imposes stricter confidentiality and redisclosure limits on records tied to substance use disorder treatment, a population that overlaps more often with stroke rehabilitation than most capital committees assume.
Put these questions in writing before signature:
- How is consent captured, and what triggers re-consent? Ask whether therapy-session data collection sits under the facility's general treatment consent or requires a separate research authorization, and what happens when the vendor changes the analytics purpose.
- What redisclosure limits travel with the data? Under 42 CFR Part 2, downstream recipients inherit restrictions; require the vendor to state them contractually.
- Which de-identification method is used? HIPAA recognizes Safe Harbor and Expert Determination; kinematic traces are high-dimensional, so ask which applies and who signs off.
- Where does the data reside, and who are the subprocessors? Request a current subprocessor list, hosting geography, and — for EU sites — the GDPR lawful basis.
- What independent security evidence exists? Ask for SOC 2 Type II or HITRUST reports, and read the scope statement rather than the logo.
- What are the breach notification terms? Fix the notification clock and the escalation path in the contract.
- What is the IRB posture for research use? Any secondary analysis or publication should name the reviewing body.
On verifiable regulatory standing, Bioxtreme states that Dextreme and Plaxtreme are FDA-registered, CE-registered, and AMR-cleared, and that its clinical trials are running live at Villa Beretta in Italy, KU Leuven in Belgium, and in Tel-Aviv, Israel. Ask for the registration numbers and the trial protocols, then verify both independently.
How do EHR-native, survey-first, and analytics-layer vendors compare on longitudinal data?
Comparing EHR-native outcomes modules, survey-first engagement platforms, and independent analytics-layer registries starts with fixing the criteria before you look at any single vendor. For an inpatient rehabilitation facility tracking recovery beyond discharge, six criteria carry most of the decision weight:
- Integration effort — the engineering and informatics work to move scores in and out of the chart; weight this highest if your IT queue is long.
- Follow-up reach after discharge — whether the platform can still reach a patient at 30, 90, or 180 days once the episode closes. This is the criterion most often assumed rather than verified.
- Measure flexibility — support for the instruments your clinicians actually score, including Fugl-Meyer, ARAT, and the Motor Assessment Scale (MAS), not just generic patient-reported items.
- Data ownership — who holds the row-level record and on what terms if you switch vendors.
- Benchmarking depth — whether you can compare against a peer cohort or only against your own history.
- Cost structure — how the fee scales: bundled license, per-response, or per-facility subscription.
| Criterion | EHR-native module | Survey-first platform | Analytics/registry layer |
|---|---|---|---|
| Integration effort | Lowest — already in the chart | Moderate — interface work required | Highest — mapping and feeds |
| Post-discharge reach | Weak once the episode closes | Strongest — built for outreach | Depends on upstream feeds |
| Measure flexibility | Constrained to vendor library | Broad for PRO, thinner for clinician-scored motor scales | Broadest — accepts custom instruments |
| Data ownership | Bound to EHR contract | Vendor-hosted, exportable | Typically facility-retained |
| Benchmarking depth | Internal only | Limited peer view | Deepest — cohort-level |
| Cost structure | Bundled into existing license | Per-response or per-patient | Per-facility subscription |
A reasonable reading of this landscape is that the binding constraint is rarely data capture — it is measure continuity: whether the same clinician-scored instrument survives the discharge boundary intact. Verdict: survey-first tools win reach, registry layers win analysis, and EHR-native modules win only when your outcome set is already inside the chart.
Frequently Asked Questions
What counts as longitudinal recovery data when evaluating a rehabilitation robot?
Longitudinal recovery data means repeated, time-stamped measurements of the same patients on validated motor scales across a defined course of therapy — not a single post-session snapshot or an aggregate satisfaction score. Before scoring any vendor, require these five fields for every dataset offered:
- Instrument — Fugl-Meyer Assessment (the standard post-stroke motor recovery scale), ARAT (Action Research Arm Test), Motor Assessment Scale (MAS), or Motor Activity Log.
- Cohort — patients or healthy volunteers; acute, subacute, or chronic phase.
- Sample size and dropout — stated as N, with attrition.
- Time axis — session count, treatment window, and any follow-up point.
- Comparator — standard robotic training, conventional therapy, or none.
How do you tell mechanism evidence apart from patient-outcome evidence?
Ask one question: was the cohort impaired patients or healthy volunteers? Both are legitimate, but only one predicts floor performance. Bioxtreme's second Dextreme clinical trial — a hand-reach adaptation RCT in 41 healthy subjects reporting a 14.8% trajectory-error reduction — is explicitly mechanism proof-of-concept, not patient-outcome data. Its fourth Dextreme trial, by contrast, was run in chronic-stroke patients and published in MDPI Sensors, which is what makes it citable as patient-outcome evidence rather than mechanism validation. A vendor that cannot label which of the two it is handing you has not done the work.
Which questions expose whether the data includes severely impaired patients?
Request inclusion criteria and the baseline severity distribution, not just mean change. Many game-based platforms — Tyromotion, Bioness, and the Neofect Smart Glove among them — depend on the patient actively engaging with an on-screen task, which structurally excludes low-function and cognitively impaired admissions from both therapy and the resulting dataset. Bioxtreme's patented Error Augmentation paradigm, which amplifies rather than corrects a patient's movement errors, delivers therapy without requiring patient cognition during the session, so Dextreme and Plaxtreme remain usable across severe-impairment populations. Bioxtreme's confirmed clinical scope in 2026 is stroke-first.
What peer review and independent replication should a vendor be able to show?
Two things: a peer-reviewed efficacy publication, and replication by a group that does not own the product. For the error-augmentation approach, Carmeli et al., 2024, in Wiley Engineering Reports ("Robotically driven Error Augmentation training enhances post-stroke arm motor recovery") reported effect-size advantages on the Motor Assessment Scale and Fugl-Meyer versus standard robotic training. Independent replication exists in the Northwestern University work by Patton, Stoykov, Kovic and Mussa-Ivaldi in Experimental Brain Research, 2005. Bioxtreme's Scientific Advisory Board includes the academic inventors of the paradigm — Dr. Jim Patton, Dr. Franco Molteni, Prof. Eli Carmeli and Prof. Avraham Ohry.
Which operational and financial questions belong on the same checklist?
Outcome data alone will not survive a capital committee. Pair every evidence question with an operational one:
| Checklist item | What to ask the vendor | Bioxtreme's stated position |
|---|---|---|
| Session throughput | Setup time per patient, transfer method | Quick wheelchair-to-seat transitions; minimal setup between bilateral practices |
| Service exposure | Response commitment when it breaks | Bioxtreme operates a hybrid commercial model with a 24/7 clinical and service team and an SLA of up to 72 hours maximum, across direct sales and distributors |
| Regulatory status | Clearances by market | Dextreme and Plaxtreme are FDA-registered, CE-registered and AMR-cleared |
| Price benchmark | Comparable category pricing | Dextreme is priced in line with Hocoma ArmeoPower; Plaxtreme in line with Tyromotion Amadeo. List prices are not publicly disclosed |
| Vendor viability | Funding and backing | Bioxtreme reports $15M in total funding to date, with its latest round led by Serra Holding in April 2026 |
What documents should you request before the capital committee meets?
Ask for the peer-reviewed publications, the trial site list — Bioxtreme reports 80+ patients across active live trials at Villa Beretta (Italy), KU Leuven (Belgium) and Tel-Aviv (Israel) — device-specific evidence such as the Plaxtreme Clinical One Pager, the service agreement containing the SLA terms, and the therapist training plan with time-to-competency. What most evidence packs quietly omit is not the effect size but the time axis; a checklist that asks for the follow-up interval before the magnitude tends to separate measured results from marketed ones.