Skip to main content

Key Takeaways:

  • Mental Workload is Dynamic: It is not simply about “focus” or “memory.” It is the constant negotiation between the demands of a task and the cognitive resources available to the specific human.
  • The Flaw of Retrospection: Post-task questionnaires can miss important real-time fluctuations in workload. Humans are highly biased when recalling difficulty, often failing to pinpoint exactly when a task was manageable or became overwhelming.
  • The Real-Time Solution: The InnoBrain Mental Workload Metric uses continuous EEG data to objectively track cognitive demand exactly when it happens, powered by InnoBrain CORTEX-O2, our latest scalable AI model.

To understand human performance, mental workload must be measured. It helps us identify if a person is dangerously overwhelmed or, equally important, inactive and losing engagement. But what exactly is mental workload?

Mental workload is often misunderstood. It is not just focus, and it is not just memory. Rather, it is a measure of how much cognitive resource an individual must spend to successfully complete a task. Crucially, workload does not depend solely on the task itself; it is a complex interaction between the objective difficulty of the work and the subjective capacity of the human at that exact moment. (For more info, see our article: Understanding Mental Workload: The Hidden Side of Human Performance)

Why Real-Time Measurement is Non-Negotiable

Historically, subjective questionnaires (such as the NASA Task Load Index (NASA-TLX) (Hart & Staveland, 1988)) have been administered at the end of tasks to assess mental workload; however, this approach has been associated with several methodological limitations.

When a human is asked to estimate how hard a task was, the data can be affected by recall bias, perception, and timing effects (Cain, 2007). More importantly, real-world tasks are not uniformly difficult (Osia et al., 2025). A person might coast through 90% of a scenario, only to experience a severe cognitive bottleneck in the 10%. If an evaluation is requested only at the end, the data tend to be averaged out, which can obscure critical fluctuations and transient peaks in workload over time.

Additionally, when the difficulty of a task is assessed at its conclusion, the placement of that 10% “high mental workload” becomes important. If the most difficult segment occurs at the very end of the task, the “recency effect” may lead to the entire experience being rated as significantly more difficult than it actually was (Hsu et al., 2018; Peterson and Kozhokar, 2017). Conversely, if the difficulty occurs early on, it may be forgotten entirely by the time the evaluation occurs. 

In addition, a distinct psychological shift can occur once a goal is reached. When people finish a task, they often view it through a lens of success, making the workload seem lighter in hindsight than it was in the heat of the moment. This “completion glow” masks the physiological and mental strain experienced during the process and biases the resulting data (Finn and Miele,2016). 

We suggest a fundamentally new method to address these problems: an objective, real-time metric derived straight from neural activity. In this article, we present InnoBrain CORTEX-O2, our newest mental workload measurement model, along with our validation test findings.

The evolution of InnoBrain CORTEX models

The development of InnoBrain’s mental workload models has followed a clear progression: from literature-based signal interpretation to scalable AI-driven workload estimation.

  • InnoBrain CORTEX-B (The Basic Model): InnoBrain CORTEX-B was our first structured mental workload model. It was built on established EEG literature and translated known neural workload markers into a computational metric (Gevins et al., 1997; Smith et al., 1999; Jaquess et al., 2018; Di Flumeri et al., 2018; Borghini et al., 2014; Smith & Gevins, 2005; Brookings et al., 1996; Wilson, 2002). This model provided a strong foundation for interpreting workload-related brain activity, but like many formula-based approaches, it had limited flexibility when dealing with real-world variation across people, tasks, and EEG hardware.
  • InnoBrain CORTEX-O (The AI Model): Building upon the foundation of InnoBrain CORTEX-B, this model introduced a learning-based approach. By leveraging machine learning, InnoBrain CORTEX-O achieved significantly higher accuracy than its formula-based predecessor (Read more about the architecture in our paper: A Real-time Unconstrained EEG-Classifier for Mental Workload Monitoring; also see our blog post: Decoding Mental Workload: From EEG Insights to Real-World Applications).
  • InnoBrain CORTEX-O2 (The Next-Generation AI): As our flagship and most advanced version, InnoBrain CORTEX-O2 maximizes accuracy while introducing unprecedented flexibility. It is hardware-agnostic, meaning it can be integrated with virtually any EEG hardware system with only a negligible drop in precision.

Validating the model: The Experimental Paradigm 

To prove that InnoBrain CORTEX-O2 can objectively measure workload without relying on biased self-reporting, we designed a rigorously controlled study using arithmetic challenges designed to push cognitive limits. 

The Study Design:

The session consisted of six blocks. In each block, participants cycled through three distinct levels of difficulty in a randomized order. Participants had one minute to solve as many problems as possible at each difficulty level. Each problem was separated by a 2-second interval, and each difficulty tier was followed by a 15-second rest period.

To create these distinct workload tiers, we manipulated the complexity of the math:

  • Easy: Single-digit addition (e.g., 4 + 5)
  • Medium: Mixed single- and double-digit addition (e.g., 8 + 14 or 22 + 45)
  • Hard: Triple-digit addition (e.g., 348 + 512)

Did this design successfully manipulate cognitive demand?

To confirm that our experimental design successfully manipulated cognitive demand, we analyzed the behavioral data. Figure 1 shows the Mean Reaction Time (RT) and Accuracy (ACC), including standard deviations (SD), across the three difficulty levels.

Repeated-measures ANOVA revealed a strong main effect of difficulty level on Reaction Time (RT): F(2,38) = 29.79, p < 0.001, η²_G = 0.47 The effect on Accuracy (ACC) was also significant: F(2,38) = 19.55, p < 0.001, η²_G = 0.37. All pairwise comparisons (Bonferroni-corrected) were statistically significant, with large effect sizes (Cohen’s d ranging from 0.99 to 2.28), showing clear differentiation between the three levels.

Figure 1: Behavioral Validation of Cognitive Demand Across Task Difficulty Levels.

Translating Behavior to Neural State: The Performance of InnoBrain CORTEX-O2

Behavioral data only tells us that performance dropped; it doesn’t show us the physiological mechanism driving the drop. For that, we turn to the EEG data.To capture these cognitive shifts in real-time, we developed InnoBrain CORTEX-O2, our most advanced model to date.

The analysis of our results, as illustrated in Figure 2, confirms that the InnoBrain CORTEX-O2 metric effectively mirrors the escalation of task difficulty. The average workload score across all subjects rose significantly as cognitive demand increased, progressing from 0.39 at Level 1 to 0.53 at Level 2, and peaking at 0.69 for Level 3. 

Notably, the neural data recorded during the rest periods (Eyes Open), often utilized as a “baseline” state in traditional research, showed mental workload levels nearly equivalent to Level 2. This is physiologically rational; despite instructions to remain “thoughtless,” human subjects lack direct control over internal cognitive processes, leading to spontaneous internal reflection or task-related rumination (Raichle et al., 2001). This finding highlights an important limitation of treating resting eyes-open periods as a true zero-workload baseline. Even during rest, participants may engage in spontaneous thought, task anticipation, reflection, or rumination (Stark & Squire, 2001).

Figure 2: Mean and Standard Deviation of InnoBrain CORTEX-O2 Predictions Across Cognitive States.

Model performance: distinguishing low and high workload

Although the Medium condition was useful for demonstrating a graded increase in predicted workload at the group level, it was excluded from binary model evaluation because its subjective difficulty varied substantially across participants. For classification analysis, we therefore focused on the two most clearly separated conditions: Easy (Level 1) and Hard (Level 3) tasks. The Medium (Level 2) tasks were excluded because they proved too difficult for some participants and too easy for others, depending on their individual mathematical skills. As a result, Medium-level tasks did not provide clear and consistent distinctions for reliable labeling, making them unsuitable for model training and evaluation in this study.

To validate how well the model differentiates between “Easy” and “Hard” cognitive states, we analyzed its classification performance. We utilized a Leave-One-Subject-Out Cross-Validation (LOSOCV) technique to verify that these findings are applicable in practical, real-world scenarios. This process involves training the model on the data from all participants except one, who is then used for testing; this cycle is repeated to confirm the system’s capacity for generalization across different individuals. Such iterative validation demonstrates that our metric does not merely memorize specific data patterns, but is a robust, subject-independent tool capable of accurately assessing previously unseen users. The results demonstrate that InnoBrain CORTEX-O2 delivers highly reliable and objective insights without the inherent biases of subjective reporting:

  • Area Under the ROC Curve (ROC AUC): This metric measures how well the model can tell the difference between a ‘low mental workload’ state and a ‘high mental workload’ state. It shows separability and measures how well the model separates the two classes, independent of a fixed threshold. A value of 0.81 signifies strong discriminative power, well above the 0.5 baseline of random classification.
    • F1-Score: This metric shows detection reliability and the balance between precision and recall. It is useful when false positives and false negatives both matter. An F1-score of 0.85 demonstrates that InnoBrain CORTEX-O2 maintains an excellent balance between minimizing false positives and capturing true high-workload instances, making it highly reliable for practical applications.
  • Balanced Accuracy: This metric shows fair performance across different classes in case there is class imbalance, when one class has more samples than the other. A score of 0.82 indicates that InnoBrain CORTEX-O2 correctly classifies mental workload levels with 82% accuracy when both classes are given equal importance. This reflects robust and consistent performance across different workload conditions.

Figure 3: Classification Performance Metrics of the InnoBrain CORTEX-O2 Model.

Tracking workload over time

The real value of a real-time workload metric becomes especially clear when looking at one participant over time.

In Figure 4, InnoBrain CORTEX-O2 predictions are shown continuously across six experimental blocks. The shaded regions indicate the different workload levels. The model output follows the changing structure of the task, rising during more demanding periods and decreasing during easier periods. This demonstrates the ability of InnoBrain CORTEX-O2 to capture dynamic workload fluctuations rather than simply assigning one general score after the task is complete. 

This is exactly where real-time EEG-based monitoring becomes powerful. Instead of asking, “How hard was the task overall?” we can begin asking:

  • When did workload increase?
  • How long did the overload last?
  • Did the person recover?
  • Did the system demand more than the person could handle at that moment?

These are the questions that matter in real-world human performance.

Figure 4: Real-Time InnoBrain CORTEX-O2 Prediction for a Single Subject Over Time.

The Future of Neuroergonomics and Human Performance

The introduction of the InnoBrain CORTEX-O2 model marks a paradigm shift in how we understand and evaluate human performance. Whether in aviation, automotive design, heavy industry, or high-stakes control rooms, operators can no longer afford to rely on flawed retrospective surveys to ensure safety and efficiency. Real-time, AI-driven EEG metrics provide an unfiltered window into the operator’s mind, allowing systems to adapt to human cognitive bottlenecks before they lead to critical errors.

References:

  1. Borghini, G., Astolfi, L., Vecchiato, G., Mattia, D., and Babiloni, F., 2014. Measuring neurophysiological signals in aircraft pilots and car drivers for the assessment of mental workload, fatigue and drowsiness. Neuroscience & Biobehavioral Reviews, 44, 58-75.
  2. Brookings, J. B., Wilson, G. F., and Swain, C. R., 1996. Psychophysiological responses to changes in workload during simulated air traffic control. Biological psychology, 42(3), 361-377.
  3. Cain, B. (2007). A review of the mental workload literature.
  4. Di Flumeri, G., Borghini, G., Aricò, P., Sciaraffa, N., Lanzi, P., Pozzi, S., … and Babiloni, F., 2018. EEG-based mental workload neurometric to evaluate the impact of different traffic and road conditions in real driving settings. Frontiers in human neuroscience, 12, 509
  5. Finn, B., & Miele, D. B. (2016). Hitting a high note on math tests: Remembered success influences test preferences. Journal of Experimental Psychology: Learning, Memory, and Cognition, 42(1), 17.
  6. Gevins, A., Smith, M. E., McEvoy, L., and Yu, D., 1997. Highresolution EEG mapping of cortical activation related to working memory: effects of task difficulty, type of processing, and practice. Cerebral cortex (New York, NY: 1991), 7(4), 374-385.
  7. Hart, S.G. and Staveland, L.E., 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology (Vol. 52, pp. 139-183). North-Holland. 
  8. Hsu, C. F., Propp, L., Panetta, L., Martin, S., Dentakos, S., Toplak, M. E., & Eastwood, J. D. (2018). Mental effort and discomfort: Testing the peak-end effect during a cognitively demanding task. PloS one, 13(2), e0191479.
  9. Jaquess, K. J., Lo, L. C., Oh, H., Lu, C., Ginsberg, A., Tan, Y. Y., … and Gentili, R. J., 2018. Changes in mental workload and motor performance throughout multiple practice sessions under various levels of task difficulty. Neuroscience, 393, 305-318.
  10. Osia, A., Tahamtan, Z., Zhao, L., Davari, M., & Nybacka, M. (2025). A Real-time Unconstrained EEG-Classifier for Mental Workload Monitoring.
  11. Peterson, D. A., & Kozhokar, D. (2017, September). Peak-end effects for subjective mental workload ratings. In Proceedings of the human factors and ergonomics society annual meeting (Vol. 61, No. 1, pp. 2052-2056). Sage CA: Los Angeles, CA: SAGE Publications.
  12. Raichle, M. E., MacLeod, A. M., Snyder, A. Z., Powers, W. J., Gusnard, D. A., & Shulman, G. L. (2001). A default mode of brain function. Proceedings of the National Academy of Sciences, 98(2), 676-682.
  13. Smith, M. E. and Gevins, A., 2005. Neurophysiologic monitoring of mental workload and fatigue during operation of a flight simulator. In Biomonitoring for Physiological and Cognitive Performance during Military Operations (Vol. 5797, pp. 116-126). SPIE.
  14. Smith, M. E., McEvoy, L. K., and Gevins, A., 1999. Neurophysiological indices of strategy development and skill acquisition. Cognitive Brain Research, 7(3), 389-404.
  15. Stark, C. E., & Squire, L. R. (2001). When zero is not zero: the problem of ambiguous baseline conditions in fMRI. Proceedings of the National Academy of Sciences, 98(22), 12760-12766.
  16. Wilson, G. F., 2002. An analysis of mental workload in pilots during flight using multiple psychophysiological measures. The International Journal of Aviation Psychology, 12(1), 3- 18.