Changing stroke rehab and research worldwide now.Time is Brain! trillions and trillions of neurons that DIE each day because there are NO effective hyperacute therapies besides tPA(only 12% effective). I have 523 posts on hyperacute therapy, enough for researchers to spend decades proving them out. These are my personal ideas and blog on stroke rehabilitation and stroke research. Do not attempt any of these without checking with your medical provider. Unless you join me in agitating, when you need these therapies they won't be there.

What this blog is for:

My blog is not to help survivors recover, it is to have the 10 million yearly stroke survivors light fires underneath their doctors, stroke hospitals and stroke researchers to get stroke solved. 100% recovery. The stroke medical world is completely failing at that goal, they don't even have it as a goal. Shortly after getting out of the hospital and getting NO information on the process or protocols of stroke rehabilitation and recovery I started searching on the internet and found that no other survivor received useful information. This is an attempt to cover all stroke rehabilitation information that should be readily available to survivors so they can talk with informed knowledge to their medical staff. It lays out what needs to be done to get stroke survivors closer to 100% recovery. It's quite disgusting that this information is not available from every stroke association and doctors group.

Showing posts with label inter-rater reliability. Show all posts
Showing posts with label inter-rater reliability. Show all posts

Wednesday, November 26, 2025

Reliability and predictive validity of the bedside upper limb functional evaluation tool (BUFET) among stroke subjects – a psychometric study

 

Predictions don't get you recovered! What are the EXACT PROTOCOLS THAT DELIVER RECOVERY?  If you can't do proper stroke research; get the hell out and do something simpler like basket weaving! You're all fired!

You can see for yourself that nothing in this Wolf Motor Test actually gets you recovered.  To me this type of testing is useless except you'll have to consent since it probably is needed to get insurance to pay. To me it would be much more useful to spend my time doing protocol repetitions leading to recovery than this shit. One question to determine patient recovery; 'Are you fully recovered? Y/N?' Then the doctor and therapist should provide EXACT PROTOCOLS THAT DELIVER RECOVERY! Oh, your incompetent? doctor doesn't have those protocols? Then you're screwed, but your doctor still gets paid for incompetence!

Wolf Motor Function Test (WMFT)

The latest here: 

Reliability and predictive validity of the bedside upper limb functional evaluation tool (BUFET) among stroke subjects – a psychometric study


We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.

Abstract

The Bedside Upper Limb Functional Evaluation Tool (BUFET) is a newly developed tool that assesses upper extremity functions, including proximal and distal movements such as gestures, gripping, and independent finger movements. This study aimed to evaluate the intra-rater and inter-rater reliability and predictive validity of the BUFET. Subjects ≥ 18 years of age and with a first episode of supratentorial stroke were included. The principal investigator, a physiotherapist, rated the subjects while performing the scale which was videotaped. To establish intra-rater reliability, the same rater re-scored the videos on two separate occasions, after 7 and 30 days. For inter-rater reliability, the principal investigator and 4 additional raters independently scored the patient based on the video. For predictive validity, the Wolf Motor Function Test (WMFT) was administered on the same day as BUFET, followed by administering the scales after 30 days. Intra- and inter-rater reliability was found to be excellent (ICC—0.992, 0.985). The reliability of all the components was > 0.89. The predictive validity was demonstrated by R2 values were 0.479 when compared with the WMFT scale. The study concluded that BUFET is an instrument with high intra-rater and inter-rater reliability with moderate predictive validity.

Wednesday, November 5, 2025

Innovative stimulated muscle contraction signals based digital muscle marker: a reliable tool for assessing muscle in persons with stroke

'Assessments' DO NOTHING unless you map EXACT RECOVERY PROTOCOLS TO THEM! This was absolutely useless for getting survivors recovered! Stroke research is to get survivors recovered; you'll want recovery when you are the 1 in 4 per WHO that has a stroke! Just maybe you want to do the proper research now!

WITH NO LEADERSHIP IN STROKE NOTHING EVER GETS DONE PROPERLY!

Inter-rater reliability and 'assessments' don't get survivors recovered; or are you that blitheringly stupid?

Innovative stimulated muscle contraction signals based digital muscle marker: a reliable tool for assessing muscle in persons with stroke


Abstract

Background

Stroke leads to motor dysfunction as a result of damage to the central nervous system, often resulting in secondary muscle changes from disuse and immobilization. Existing tools for assessing these muscle changes in persons with stroke are limited, as these methods rely on voluntary muscle contractions, which are frequently impaired in hemiparesis. Additionally, these assessments often require complex and cumbersome equipment. Therefore, a novel system that can address these issues, with high levels of reliability is needed. This study introduces and evaluates the reliability levels of a novel technique using stimulated muscle contraction signals (SMCS)-based digital muscle markers.

Methods

SMCS-based digital muscle markers were measured on both thighs of participants with hemiplegic stroke (n = 40). Three trials were conducted by two examiners using the exoPill device, which captures muscle contraction signals induced by electrical stimulation. Isometric knee extensor strength and hand grip strength were measured, and appendicular lean mass was estimated using bioelectrical impedance analysis. The intra- and inter-rater reliability of SMCS-based digital muscle marker measurements was evaluated using the intraclass correlation coefficient (ICC), coefficient of variance (CV), and 95% limits of agreement (LOA).

Results

The ICC for both intra- and inter-rater reliability of SMCS-based digital muscle marker measurements exceeded 0.98 on both the affected and unaffected sides. The CV was consistently low, ranging from 1.36% to 3.07% in the unaffected leg, and remained under 3% in the affected leg. Bland-Altman plots confirmed consistent agreement, with most measurements falling within the 95% LOA and no systematic bias were observed. Mean differences for intra- and inter-rater assessments ranged from − 0.43 to 0.25 for the unaffected leg and from 0.00 to 0.05 for the affected leg. Spearman correlation coefficient showed moderate to good correlation (ρ = 0.50) between SMCS-based digital muscle markers and knee extensor strength for both unaffected and affected leg, and good to excellent correlations (ρ = 0.87) between SMCS-based digital muscle markers and leg lean mass.

Conclusions

SMCS-based digital muscle markers demonstrated high reliability and significant correlations with muscle strength and mass, suggesting their potential utility for assessing muscle in individuals with stroke.

Thursday, October 23, 2025

The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation-based bimanual upper-limb performance measurement

 

Assessments DO NOTHING unless you map EXACT RECOVERY PROTOCOLS TO THEM! This was absolutely useless, NOTHING ON PROTOCOLS THAT WILL DELIVER RECOVERY!  

Inter-rater reliability does nothing for survivor recovery!

The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation-based bimanual upper-limb performance measurement

Objective: To investigate interrater and intrarater reliability, measurement error and convergent and discriminative validity of the Adult Assisting Hand Assessment Stroke (Ad-AHA Stroke).

Design: Cross-sectional observational study

Setting: Seven stroke rehabilitation centers

Participants: A total of 118 stroke survivors (reliability sample: n=30; validity sample: n=118) were included (median age 67 years (interquartile range (IQR) 59-76); median time post stroke 81 days (IQR 57-117).

Interventions: N/A.

Main Outcome Measures: Ad-AHA Stroke, Action Research Arm Test (ARAT), Upper Extremity Fugl-Meyer assessment (UE-FMA). The Ad-AHA Stroke is an observation-based instrument assessing the effectiveness of the spontaneous use of the affected hand when performing bimanual activities in adults after stroke. Reliability of Ad-AHA stroke was examined using intraclass correlation coefficients (ICC), Bland-Altman plots, and weighted kappa (Kw) statistics for reliability on item level. Standard error of measurement (SEM) was calculated based on Ad-AHA units. Convergent validity was assessed by calculating Spearman rank correlation coefficients between Ad-AHA stroke and ARAT and UE-FMA. Comparison of Ad-AHA stroke scores between subgroups of patients according to hand dominance, neglect and age evaluated discriminative validity.

Results: Intrarater and interrater agreement showed an ICC of 0.99 (95% CI=0.99-0.99), a SEM of 2.15 and 1.64 out of 100, respectively and Kw for item scores were all above 0.79. The relation between Ad-AHA and other clinical assessments was strong (rs=0.9). Patients with neglect had significantly lower Ad-AHA scores compared to patients without neglect (p=0.004).

Conclusion: The Ad-AHA Stroke captures actual bimanual performance. Thereby it provides an additional aspect of upper limb assessment with good to excellent reliability and low SEM for patients with sub-acute stroke. High convergent validity with ARAT and UE-FMA and discriminative validity was demonstrated.


Wednesday, September 24, 2025

The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation-based bimanual upper-limb performance measurement

Absolutely useless, NOTHING ON PROTOCOLS THAT WILL DELIVER RECOVERY! Assessments DO NOTHING! Interrater crapola DOES NOTHING TO GET SURVIVORS RECOVERED! Does no one in stroke know how to think? I'd have you all fired for incompetence!
The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation-based bimanual upper-limb performance measurement

The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation- 
based bimanual upper-limb performance measurement 
Objective: To investigate interrater and intrarater reliability, measurement error and 
convergent and discriminative validity of the Adult Assisting Hand Assessment Stroke (Ad- 
AHA Stroke). 
Design: Cross-sectional observational study 
Setting: Seven stroke rehabilitation centers 
Participants: A total of 118 stroke survivors (reliability sample: n=30; validity sample: n=118) 
were included (median age 67 years (interquartile range (IQR) 59-76); median time post 
stroke 81 days (IQR 57-117). 
Interventions: N/A. (With no interventions, this was absolutely fucking useless for survivors!)
Main Outcome Measures: Ad-AHA Stroke, Action Research Arm Test (ARAT), Upper 
Extremity Fugl-Meyer assessment (UE-FMA). The Ad-AHA Stroke is an observation-based 
instrument assessing the effectiveness of the spontaneous use of the affected hand when 
performing bimanual activities in adults aſter stroke. Reliability of Ad-AHA stroke was 
examined using intraclass correlation coefficients (ICC), Bland-Altman plots, and weighted 
kappa (Kw) statistics for reliability on item level. Standard error of measurement (SEM) was 
calculated based on Ad-AHA units. Convergent validity was assessed by calculating Spearman 
rank correlation coefficients between Ad-AHA stroke and ARAT and UE-FMA. Comparison of 
Ad-AHA stroke scores between subgroups of patients according to hand dominance, neglect 
and age evaluated discriminative validity. 
Results: Intrarater and interrater agreement showed an ICC of 0.99 (95% CI=0.99-0.99), a 
SEM of 2.15 and 1.64 out of 100, respectively and Kw for item scores were all above 0.79. 

Wednesday, August 27, 2025

The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation-based bimanual upper-limb performance measurement

I see zero use for assessments. Why not deliver EXACT STROKE PROTOCOLS THAT DELIVER 100% RECOVERY,  instead of this lazy shit. Inter-rater reliability does nothing for survivor recovery!

Oops, I'm not playing by the polite rules of Dale Carnegie,  'How to Win Friends and Influence People'. 

Telling stroke medical persons they know nothing about stroke is a no-no even if it is true. 

Politeness will never solve anything in stroke. Yes, I'm a bomb thrower and proud of it. Someday a stroke 'leader' will try to ream me out for making them look bad by being truthful , I look forward to that day. 

(Where is the creation of protocols that deliver recovery? Without that this research is useless!)
The Adult Assisting Hand Assessment Stroke: Psychometric properties of an observation-based bimanual upper-limb performance measurement

Van Gils A, Meyer S, Van Dijk M, Thijs L, Michielsen M, Lafosse C, Truyens V, Oostra
K, Peeters, A, Thijs V, Feys H, Krumlinde-Sundholm L, Kos D , Verheyden G
Abstract
Objective: To investigate interrater and intrarater reliability, measurement error and convergent
and discriminative validity of the Adult Assisting Hand Assessment Stroke (Ad-AHA Stroke).
Design: Cross-sectional observational study
Setting: Seven stroke rehabilitation centers
Participants: A total of 118 stroke survivors (reliability sample: n=30; validity sample: n=118) were
included (median age 67 years (interquartile range (IQR) 59-76); median time post stroke 81 days
(IQR 57-117).
Main Outcome Measures: Ad-AHA Stroke, Action Research Arm Test (ARAT), Upper Extremity
Fugl-Meyer Assessment (UE-FMA). The Ad-AHA Stroke is an observation-based instrument
assessing the effectiveness of the spontaneous use of the affected hand when performing bimanual
activities in adults after stroke. Reliability of Ad-AHA stroke was examined using intraclass
correlation coefficients (ICC), Bland-Altman plots, and weighted kappa (Kw) statistics for reliability
on item level. Standard error of measurement (SEM) was calculated based on Ad-AHA units.
Convergent validity was assessed by calculating Spearman rank correlation coefficients between
Ad-AHA stroke and ARAT and UE-FMA. Comparison of Ad-AHA stroke scores between subgroups
of patients according to hand dominance, neglect and age evaluated discriminative validity.
Results: Intrarater and interrater agreement showed an ICC of 0.99 (95% CI=0.99-0.99), a SEM of
2.15 and 1.64 out of 100, respectively and Kw for item scores were all above 0.79. The relation
between Ad-AHA and other clinical assessments was strong (ρ=0.9). Patients with neglect had
significantly lower Ad-AHA scores compared to patients without neglect (ρ=0.004).
Conclusion: The Ad-AHA Stroke captures actual bimanual performance. Thereby it provides an
additional aspect of upper limb assessment with good to excellent reliability and low SEM for
patients with sub-acute stroke. High convergent validity with ARAT and UE-FMA and discriminative
validity was supported.

Monday, August 11, 2025

Inter-rater reliability and validity of the stroke rehabilitation assessment of movement (STREAM) instrument

 But you failed at the next step. You didn't create EXACT RECOVERY PROTOCOLS based upon your assessment! 

Inter-rater reliability is useless because it means you're doing a subjective diagnosis, NOT AN OBJECTIVE DIAGNOSIS! Since you don't understand what you're doing: YOU'RE FIRED! along with your mentors and senior researchers! Stroke research is meant to get survivors recovered! THIS COMPETELY FAILED AT THAT!

Inter-rater reliability and validity of the stroke rehabilitation assessment of movement (STREAM) instrumentr reliability and validity of the stroke rehabilitation assessment of movement (STREAM) instrument

Chun-Hou Wang , Ching-Lin Hsieh , May-Hui Dai , Chia-Hui Chen , Yu-Fen Lai
DOI: 10.1080/165019702317242668

Abstract

The Stroke Rehabilitation Assessment of Movement (STREAM) instrument is used to measure motor and mobility problems in patients who have experienced a stroke. The purposes of the study were to examine the inter-rater reliability, concurrent and convergent validity of the STREAM instrument in stroke patients. Fifty-four stroke patients participated in the study. For the purpose of inter-rater reliability, the STREAM instrument was administered by two raters on each patient within a 2-day period. Validity was assessed by comparing the patients' scores on the STREAM instrument with those obtained from the other well-established measures. Weighted kappa statistics for inter-rater agreement on scores for individual items ranged from 0. 55 to 0. 94. The intraclass correlation coefficient for the total score was 0. 96 indicating very high inter-rater reliability. The intraclass correlation coefficients were also very high in each of the subscales. The total STREAM score was moderately to highly associated with the score of the Barthel Index and Fugl-Meyer motor assessment scale, rho = 0. 67, and 0. 95, respectively. The STREAM subscale scores were closely associated with scores of the other well-validated measures. Our results demonstrate that consistent and valid information can be obtained from the STREAM instrument and support its use in the value of the STREAM evaluation of motor and mobility recovery in persons who have experienced a stroke.

Lay Abstract

Wednesday, July 2, 2025

The Reliability of the Wolf Motor Function Test for Assessing Upper Extremity Function After Stroke

 

You can see for yourself that nothing in this Wolf Motor Test actually gets you recovered.  To me this type of testing is useless except you'll have to consent since it probably is needed to get insurance to pay. To me it would be much more useful to spend my time doing protocol repetitions leading to recovery than this shit. Please talk to survivors sometime and see what they want out of research and stroke rehab. Not this crapola!  'Assessments' do nothing for recovery!


Wolf Motor Function Test (WMFT)

The latest here:

The Reliability of the Wolf Motor Function Test for Assessing Upper Extremity Function After Stroke

Morris DM, Uswatte G, Crago JE, Cook EW III, Taub E. 
Objective: 

To examine the reliability of the Wolf Motor Function Test (WMFT) for assessing upper extremity motor function in adults with hemiplegia. 

Design: 

Interrater and test-retest reliability. (Nothing about interrater stuff gets survivors recovered! Useless!)

Setting: 

A clinical research laboratory at a university medical center. 

Patients: 

A sample of convenience of 24 subjects with chronic hemiplegia (onset >1yr), showing moderate motor impairment. 

Intervention: 

The WMFT includes 15 functional tasks. Performances were timed and rated by using a 6-point functional ability scale. The WMFT was administered to subjects twice with a 2-week interval between administrations. All test sessions were videotaped for scoring at a later time by blinded and trained experienced therapists. 

Main Outcome Measure: 

Interrater reliability was examined by using intraclass correlation coefficients and internal consistency by using Cronbach's alpha. Results: Interrater reliability was.97 or greater for performance time and.88 or greater for functional ability. Internal consistency for test 1 was.92 for performance time and.92 for functional ability; for test 2, it was.86 for performance time and.92 for functional ability. Test-retest reliability was.90 for performance time and.95 for functional ability. Absolute scores for subjects were stable over the 2 test administrations. 
Conclusion: 

The WMFT is an instrument with high interrater reliability, internal consistency, test-retest reliability, and adequate stability. (But DOES ABSOLUTELY NOTHING TOWARDS RECOVERY!)

Tuesday, June 3, 2025

Inter-rater reliability and validity of the stroke rehabilitation assessment of movement (STREAM) instrument

Assessments do nothing unless they are directly followed by EXACT PROTOCOLS THAT DELIVER RECOVERY!  Inter-rater reliability is useless because it means you're doing a subjective diagnosis, NOT AN OBJECTIVE DIAGNOSIS! Since you don't understand what you're doing: YOU'RE FIRED!

Inter-rater reliability and validity of the stroke rehabilitation assessment of movement (STREAM) instrument

Chun-Hou Wang, 1 Ching-Lin Hsieh, 2 May-Hui Dai, 3 Chia-Hui Chen 3 and Yu-Fen Lai 3 From the 1 School of Rehabilitation, Chung-Shan Medical and Dental College, Taichung, 2 School of Occupational Therapy, College of Medicine, National Taiwan University, Taipei, and 3 Department of Physical Therapy, Chung-Shan Medical and Dental College Hospital, Taichung, Taiwan, ROC 

The Stroke Rehabilitation Assessment of Movement (STREAM) instrument is used to measure motor and mobility problems in patients who have experienced a stroke. The purposes of the study were to examine the inter- rater reliability, concurrent and convergent validity of the STREAM instrument in stroke patients. Fifty-four stroke patients participated in the study. For the purpose of inter- rater reliability, the STREAM instrument was administered by two raters on each patient within a 2-day period. Validity was assessed by comparing the patients’ scores on the STREAM instrument with those obtained from the other well-established measures. Weighted kappa statistics for inter-rater agreement on scores for individual items ranged from 0.55 to 0.94. The intraclass correlation coefficient for the total score was 0.96 indicating very high inter-rater reliability. The intraclass correlation coefficients were also very high in each of the subscales. The total STREAM score was moderately to highly associated with the score of the Barthel Index and Fugl-Meyer motor assessment scale, rho = 0.67, and 0.95, respectively. The STREAM subscale scores were closely associated with scores of the other well- validated measures. Our results demonstrate that consistent and valid information can be obtained from the STREAM instrument and support its use in the value of the STREAM evaluation of motor and mobility recovery in persons who have experienced a stroke. 

Key words: cerebrovascular accident, reproducibility of results, movement. 

J Rehabil Med 2002; 34: 20–24 Correspondence address: Dr Ching-Lin Hsieh, 7 Chun- Shan South Road, School of Occupational Therapy, College of Medicine, National Taiwan University, Taipei 100, Taiwan, ROC. E-mail: mike26@ha.mc.ntu.edu.tw 

Sunday, December 22, 2024

Evaluating inter- and intra-rater reliability in assessing upper limb compensatory movements post-stroke: creating a ground truth through video analysis?

 Crapola like this DOES NOTHING to get survivors recovered!

Who gives a shit about inter-rater reliability you blithering idiots? Certainly not survivors!

Evaluating inter- and intra-rater reliability in assessing upper limb compensatory movements post-stroke: creating a ground truth through video analysis?

Abstract

Background

Compensatory movements frequently emerge in the process of motor recovery after a stroke. Given their potential for unfavorable long-term effects, it is crucial to assess and document compensatory movements throughout rehabilitation. However, clinically applicable assessment tools are currently limited. Deep learning methods have shown promising potential for assessing movement quality and addressing this gap. A crucial prerequisite for developing an accurate measurement tool is ensuring reliability in assessing compensatory movements, which is essential for establishing a valid ground truth.

Objective

The study aimed to assess inter- and intra-rater reliability of occupational and physical therapists’ visual assessment of compensatory movements based on video analysis.

Methods

Experienced therapists evaluated video-recorded performances of a standardized drinking task through an online labeling system. The standardized drinking task was performed by seven individuals with mild to moderate upper limb motor impairments after a stroke. The therapists rated compensatory movements in predetermined body segments and movement phases using a slider with a continuous scale ranging from 0 (no compensation) to 100 (maximum compensation). The collected data were analyzed using a generalized-linear mixed effects model with zero-inflated beta regression to estimate variance components. Intraclass correlation coefficients (ICC) were calculated to assess inter- and intra-rater reliability.

Results

Twenty-two therapists participated in this study. Inter-rater reliability was good for the phases of reaching, drinking, and returning (ICC ≥ .0.75), and moderate for both phases of transporting. Intra-rater reliability was excellent for the drinking phase (ICC > 0.9) and moderate to good for the phases of reaching, transporting, and returning of our cohort. ICCs for smoothness and interjoint coordination were poor for both inter- and intra-rater reliability. The data analysis unveiled a wide range of credible intervals for the ICCs across all domains examined in this study.

Conclusions

While this study shows promising inter- and intra-rater reliability for the drinking phases within our sample, the wide credible intervals raise the possibility that these results may have occurred by chance. Consequently, we cannot recommend the establishment of a ground truth for the automatic assessment of compensatory movements during a drinking task based on therapists’ ratings alone.

Background

The Global Burden of Disease Study 2019 shows an increase in overall stroke incidents and exposes stroke as the third-leading cause of death and disability combined. About 80% of stroke survivors are affected by motor deficits, with approximately 50% experiencing persistent upper limb impairments [17, 27, 42].

Considering the involvement of arm and hand functions in activities of daily living (ADL), facilitating upper limb motor recovery following a stroke is essential. According to the current state of scientific consensus, motor recovery is distinguished by true recovery and motor compensation [8, 24, 32]. The Stroke Recovery and Rehabilitation Roundtable (SRRR) defines true recovery as regaining the same movement patterns as available before the injury and motor compensation as the development of new motor patterns by using intact body structures to accomplish an activity goal [8]. The definitions provided by Levin et al. [32] and Kleim [24] distinguish between motor compensation and true recovery across three different levels of the international Classification of Functioning, Disability and Health (ICF [62]) domains: health condition, body function/structure and activity. According to Levin et al. [32] changes at the Health Condition level occur through processes at the neuronal level. On this level, motor compensation is characterized by structural reorganization within the brain, where other brain regions assume the functions of damaged areas. At the body function/structure level, compensation is reflected in alternative movement patterns. At the activity level, compensation is evident when different limbs or end effectors are used compared to the premorbid status. Achieving a desired goal with the impaired arm can promote its use and prevent learned non-use, potentially enhancing functional capacity and independence in activities of daily living [21, 22]. However, compensatory movements can lead to musculoskeletal changes, increasing the risk of chronic pain [12, 13, 32]. Furthermore, compensation with the unimpaired limb can negatively impact neural reorganization, potentially hindering the true recovery of the impaired limb [22, 46]. Therefore, it is important to support the true recovery on the body function/structure level during rehabilitation, rather than allowing alternative movement patterns, to minimize adverse side effects. However, to facilitate the identification and treatment of motor compensation, a comprehensive assessment is necessary [10, 26]. Compensatory movements can be assessed by measuring the quality of upper limb movement patterns through motion analysis. Possible approaches to conduct this analysis include qualitative descriptions of movements based on visual observation, observer-based scoring using standardized assessment tools or kinematic motion analysis technologies, such as marker-based motion capture systems [50].

In research settings, marker-based motion capture systems are considered the gold standard for motion analysis. Unfortunately, its implementation in clinical practice is only possible to a limited extent because of high costs, duration of the assessment as well as requirement for high-quality instruments and specialized training [60].

Core measurement sets recommended to be applied in motor rehabilitation and recovery trials post-stroke describe the Fugl-Meyer Assessment (FMA, [16]) as essential tool for evaluating upper limb body functions, and the Action Research Arm Test (ARAT, [33]) for assessing motor function at the activity level [25, 41]. However, previous research indicated that these assessments simply focus on the ability to complete a task, without capturing the movement quality. Hence they do not differentiate between the recovery of movement patterns and the use of compensatory movement strategies [26, 32, 44]. In response to the need for an assessment specifically designed to measure motor compensation, Levin et al. [30] developed the Reaching Performance Scale for Stroke (RPSS). The RPSS [30] is an ordinal-scale assessment tool that evaluates compensatory movements of the upper extremity during two specific reaching tasks, excluding drinking motions. Recent studies by Subramanian et al. [54, 55] on the RPSS [30] provide promising results, describing it as a valid, reliable, and responsive scale for visually assessing motor compensation. However, the assessment is susceptible to ceiling effects due to its ordinal scaling. To our knowledge the RPSS is not yet commonly used in clinical practice. Consequently, therapists assess compensatory movements qualitatively through visual observation relying on their clinical expertise and experience [14, 28, 44, 47].

In recent years, considerable research has focused on developing assessments for motor compensation. However, due to the complexity of developing accurate and sensitive observer-based scoring assessments and the lack of consensus regarding the use of marker-based motion capture systems, further research is required to construct reliable, feasible and affordable measurement tools [26, 29, 47, 51].

Previous research has demonstrated that kinematic motion analysis technologies, such as marker-based motion capture systems, wearable sensors, marker-free vision sensors (e.g. simple cameras, or Microsoft Kinect depth sensors) or sensors embedded in rehabilitation training systems, provide objective, consistent, and precise detection of compensatory movements through automated processes, devoid of any floor and ceiling effects respectively [23, 28, 36, 60]. Recent reviews [39, 47] have highlighted the advantages of Machine Learning (ML) algorithms, which include high agreement levels, high accuracy, and low-cost, unobtrusive home based monitoring of patients. Supervised ML algorithms depend on labeled data, which serves as ground truth, to derive an optimal model capable of accurately predicting outcomes or classifying new, unseen data by leveraging learned patterns from the training dataset [5, 40]. In stroke rehabilitation, a promising domain for applying ML algorithms lies in the automated analysis of movement patterns in impaired limbs [39].

Therefore, we are currently conducting larger project to develop a ML-supported tool to analyze compensatory movements of the upper extremities and trunk on the body function/structure level, which can be used remotely with simple methods in the homes of patients after stroke [52, 56]. In this larger study, we plan to utilize simple webcam or smartphone cameras to assess compensation during a drinking task [1]. During the development of the ML algorithm, the completion of the task is recorded with a webcam and manually labeled by experienced therapists, who assess the extent of compensatory movements. This data was intended to be used to create a ground truth for the ML algorithm. While therapists are trained to assess movement quality in person (3-dimensional), it is unclear if they can do so when rating a drinking task based on videos (2-dimensional). The findings from Martinez et al. [35], Bernhardt et al. [6] and Bernhardt et al. [7] demonstrated good reliability in assessing reaching and object-lifting tasks using video analysis. However, these studies did not specifically evaluate the drinking task.

We seek to imitate human decision-making behavior using ML-supported models. To train an ML algorithm to imitate human intelligence, i.e. to build a ground truth that reflects human knowledge, we must make implicit therapeutic decisions, such as rating compensatory movements, explicitly available. This raises a broader discussion, which we intend to explore in this article. If we want ML to learn from human intelligence, a key challenge is determining whether humans can reliably assess movements during a drinking task from 2-dimensional video recordings in the first place. Depending on the findings, it remains uncertain whether basing a ground truth for assessment on human decisions would be a viable or effective approach.

Therefore, we conducted this study to investigate the inter-rater and intra-rater reliability of experienced therapists’ assessment of compensatory movement behavior based on video recordings and hence infer if such ratings are suitable for creating a ground truth for an ML application.