Summary
Most measurement of training programs stops at determining the level of "trainee satisfaction", capturing momentary impressions influenced by several factors even though training may not change performance in the workplace. Related studies and reports therefore confirm that more than half of training expenditure is wasted as "scrap learning" because measurement and application are not designed from the outset. The solution lies in a system of six pillars: backward design, distinguishing effort from impact, measurement over time, impact isolation, partnership with the direct manager, and a dual language of value (ROE/ROI). From Kaizen Consulting's perspective, measurement is born alongside training-needs analysis, not after program delivery. Impact is tracked as an extended "film" (before/during/after), not as a "snapshot" in the room. Financial return is calculated only on impact that has actually been demonstrated, while the language of value is matched to the stakeholder's question: "Return on Expectations" for the operational sponsor and "Return on Investment" for the company's finance executive.
Introduction
Most training measurement captures trainee satisfaction at a moment of heightened emotion: everyone applauds in the room, yet the effect is absent from the workplace. Global reports confirm that more than half of training expenditure is wasted because measurement and application are overlooked when the program is designed.
This article presents a practical system of six pillars that shifts the question from "Did the trainees like it?" to "What change did the training produce?" The pillars are backward design, distinguishing effort from impact, measurement over time, impact isolation, partnership with the direct manager, and a dual language of value that addresses the operational sponsor through Return on Expectations and the finance executive through Return on Investment. Together, these pillars link training effort to behavioral impact, then to organizational outcomes, and ultimately to financial and non-financial returns.
1. When Everyone Applauds and No One Changes
Why does training earn applause in the room, yet leave no visible effect in the workplace?
Imagine a training program that has just ended. The room is filled with applause; everyone gathers for a farewell photograph after an emotional speech in which the trainer apologizes for any shortcomings and lavishly thanks the trainees. The program supervisor smiles with elation, the evaluation forms are covered in stars - 4.9 out of 5 - and the report rises to the management of the beneficiary organization as conclusive proof of success. Then the weeks pass: no behavior changes at work, no indicator moves on the performance dashboard, and it is as though nothing happened. This is the paradox that haunts the training industry: everyone applauds... and no one changes. Nor is this merely a theoretical conclusion; several incidents remain lodged in memory, each revealing a different aspect of the problem.
During a visit to the social responsibility department of a financial institution, the discussion flowed as we traced the training journey from needs analysis to program completion. The moment we raised the idea of impact measurement; however, the enthusiasm froze. Displeasure appeared on the manager's face. He said tersely, "There is no need for that," then offered his justification: "Impact measurement is complex, takes a long time, and will consume an additional budget."
In a side conversation with a colleague at a training company implementing a project for a large organization, he described how they proposed impact measurement to the project owner. The response was completely beside the point: "Do not keep the trainees in the room for five hours; there are enough, including the breaks. I sent my colleagues to have a change of atmosphere."
The third incident came from a colleague who supervises training programs. Commenting on the high satisfaction reports with revealing irony, he said: "I know this trainer well. He is generous with long breaks and dismisses the trainees early. These results do not reflect learning; they merely reflect a dopamine rush, nothing more." His remark neatly condenses the entire predicament of the familiar first level of measurement: even a satisfaction form may measure the trainer's generosity with breaks rather than the value of what attendees learned.
These incidents are not rare; they may be the recurring rule, accompanied by rituals that are almost sacred. We measure attendance through signatures and fingerprints, learning by the number of hours, and success through an emotional impression collected at the very moment when the trainee is most affected and least objective. Then, when the finance director asks the question that unsettles everyone - "What return did we obtain from what we spent?" - all we have is the same satisfaction form. At that point, the painful gulf is exposed between training that delights the room and training that changes the organization.
The problem is not the absence of measurement, but its superficiality. We measure what is easy to measure, not what deserves to be measured. This article is a journey out of that predicament towards an integrated system that links training effort to behavioral impact, organizational outcomes, and financial and non-financial returns. It takes us from the question "Did they like the training?" to "What change did the training produce?" If you have ever applauded in a training room and then found no trace of it in the workplace, the following pages were written for you.
2. What Do Reports Say About Poor Measurement Practice?
When we move from individual experience to globally published data, we discover that our difficulty is not exceptional; it is the prevailing rule across the global training industry. (See the Saudi applied study on the impact of training on job performance: Al-Anazi et al., 2023.)
- Measurement stops at the first threshold: Kirkpatrick's four-level model - reaction, learning, behavior, and results - has been the world's best-known framework for evaluating training since the 1950s. The paradox is that most organizations settle for their lowest level. Various estimates indicate that approximately 80% of training events are measured at Level 1 (reaction) (Training Industry, 2026), while most organizations stop at Levels 1 and 2, entirely missing the measurement of behavioral change and its business impact (Rcademy, 2026).
- Scrap learning: Researcher Robert Brinkerhoff coined the term "scrap learning" to describe what an employee is trained to do but never applies at work. His research shows that approximately 20% of trainees never apply what they have learned, while 65% attempt to apply it and then revert to their old habits. In other words, between 80% and 85% of training is wasted and never used consistently in the workplace (Brinkerhoff & Mooney, 2008). Even the more conservative estimates offer little reassurance. In 2014, the Corporate Executive Board (CEB) estimated that 45% of all learning is not applied at work (Corporate Executive Board, 2014, as cited in Phillips, 2016). Whatever the correct figure - 45% or 85% - the result is the same: half or more of training expenditure goes to waste because measurement and application were not designed from the outset.
- Measurement tools have remained stagnant for years: Successive LinkedIn Workplace Learning reports converge on the same diagnosis: most training functions still measure success in the same old way - qualitative feedback, numbers of courses, and completion rates - metrics that do not capture the true impact of learning on business performance (LinkedIn Learning, 2022). As pressure from executive leadership grows for clear evidence of the return generated by training programs, the measurement gap has become a burden on the credibility of the learning and development function itself (LinkedIn Learning, 2025).
3. The Training Market and Its Returns:
How much do organizations around the world spend on training, and how much of that expenditure is wasted?
It is enough to consider the scale of global expenditure on corporate training. Research firms' estimates vary, but they converge on an enormous figure: the global corporate training market was valued at approximately $412.8 billion in 2024 and is expected to exceed $808 billion by 2033, at a compound annual growth rate of about 7.8% (SkyQuest Technology, 2025).
Other reports confirm the same trajectory. Allied Market Research expects the market to rise from $361.5 billion in 2023 to approximately $805.6 billion by 2035 (Allied Market Research, 2025). Variation among research firms is natural and not a cause for concern; it reflects differences in the base year, scope of definition, and methodology. Yet all estimates converge in one unmistakable direction: this is an industry whose size will double within a decade and that consumes enough of organizational budgets to deserve - indeed, require - rigorous measurement of its return.
The other face of these billions:
When the proportions of "scrap learning" are applied to this enormous expenditure, the scale of the loss becomes clear. Association for Talent Development (ATD) data indicate that average direct expenditure on learning per employee was approximately $1,229 for 32.4 training hours annually. Applying the conservative waste rate of 45% means that approximately $553 and 14.6 hours are lost per employee. Applying Brinkerhoff's higher estimate raises the waste to approximately $983 and 25.9 hours per employee (Phillips, 2016).
When the equation encompasses tens of thousands of employees across thousands of organizations, we realize that this is not a minor procedural defect but a global financial hemorrhage amounting to billions of dollars every year, driven primarily by the absence of a system that measures impact and manages application.
4. Why Does Traditional Measurement Fail?
Before proposing an alternative, it is useful to diagnose the roots of the problem which recur across contexts:
- Measurement is an afterthought: the program is designed first, and only at the end does the organization look for a way to measure it, making it impossible to link the program to objectives that were never defined in advance.
- No one asks for more: if leadership is satisfied with asking, "Did they like the training?", the learning function will measure nothing beyond that (Valamis, 2026).
- There is no infrastructure for impact measurement: tools, models, and even specialized teams are lacking.
- The role of the direct manager in developing training impact is neglected. Application in the workplace is sustained through follow-up by the direct leader, which is the decisive factor in reducing waste (Brinkerhoff, in Chief Learning Officer, 2011).
- Activity is confused with impact: the numbers of courses and hours are indicators of effort, not indicators of results.
5. Global Experiences That Overcame the Challenge:
How have leading organizations addressed the challenge of measuring training impact, and what can we learn from their experiences?
The industry did not surrender to this reality. It developed models and schools of thought, each attempting to close a gap in the measurement system. These include:
First: The Kirkpatrick School
Donald Kirkpatrick established the logic of a "chain of evidence" through four progressive levels, from learner satisfaction to business results. In its contemporary development - the "New World Kirkpatrick Model" by Jim and Wendy Kirkpatrick - emphasis was placed on beginning with the intended result (backward design) and on "Return on Expectations" (ROE) as the deepest indicator of value for stakeholders (Kirkpatrick & Kirkpatrick, 2016).
The four levels in this chain progress upwards as follows:
- Reaction (Did the trainee find the experience useful and relevant?).
- Learning (What knowledge, skills, and attitudes did the trainee acquire?).
- Behavior (Did the trainee apply what was learned in the workplace?).
- Results (What impact did the organization achieve consequently?).
It is called a "chain of evidence" because each level prepares the way for the next: satisfaction facilitates learning, learning opens the door to behavior, and behavior leads to results. If the chain breaks at any level, the levels beyond it cannot be reached. (Kachroud & Riyadh, 2020.)
The contribution of the "New World Model", however, goes beyond arranging the levels and adds practical mechanisms. It places Level 3 (behavior) at the heart of the system and surrounds it with what it calls "required drivers" - reinforcement, reminders, follow-up, and rewards - because behavior without environmental support quickly regresses.
It also introduces "leading indicators" that provide early evidence that behavior is on course to produce the desired result. The organization therefore does not wait for the impact to be complete before knowing whether it is on the right path. In this way, the model shifts from a "post-event evaluation tool" to a "performance navigation map" that begins with the result and is managed throughout the journey. (Kirkpatrick & Kirkpatrick, 2016; Kirkpatrick Partners, 2024.)
Second: Phillips Methodology
Jack Phillips added a fifth level above Kirkpatrick's four levels: Return on Investment (ROI), expressed as the ratio between program benefits converted into monetary value and the program's full cost. The methodology's most important contribution is its "isolation techniques" - control groups, trend lines, and forecasting - which separate the effect of training from other variables and make the attribution rate defensible before financial management (Phillips, 2003; Za'eemi, 2016).
Phillips extends Kirkpatrick rather than replacing it. The four levels remain the chain of evidence - reaction, learning, behavior, and results - while the fifth level translates results into the language of money through a recognized "systematic process": identify improvement in a business indicator, isolate the portion attributable to training, convert it into monetary value, and then balance it against the fully loaded cost. In this way, the return is calculated only on impact that has genuinely been demonstrated, not on an assumed promise.
Two rules protect this methodology from challenge: explicit isolation, which attributes credit to training alone rather than to the market, leadership, or incentives; and conservative estimation, which builds figures on the most reliable sources and places benefits that cannot credibly be monetized in the category of intangible benefits. The return ratio thereby shifts from a promotional figure to an argument that can withstand scrutiny from the finance director - an issue discussed in detail under the sixth pillar, on the dual language of value (Phillips, 2003).
Third: The 70-20-10 Model:
This model originated in research by the Center for Creative Leadership (CCL) in the 1980s, when approximately 191 successful executives were asked about the sources of their learning. The findings showed that the largest share came from direct work experience (70%), followed by social learning such as coaching and feedback (20%), and finally formal, structured training (10%) (McCall, Lombardo, & Morrison, 1988). Its measurement implication is exceptionally clear: if most learning occurs outside the training room, measurement must not be confined to the room.
Fourth: Corporate Universities: Crotonville as a Model
General Electric's university at Crotonville - founded in 1956 - is one of the world's earliest corporate universities. Under Jack Welch, it became the nerve center of a comprehensive cultural transformation rather than a place for delivering courses. Its measurement lies in linking leadership development directly to strategic business transformations, so that training became a lever for organizational change measured by its results, not its outputs (IMD, 2025).
This was evident in its practical mechanism. Crotonville linked its programs to an "action learning" approach: participants worked on real strategic problems facing the company, then presented their solutions to leadership, which adopted those suitable for implementation. Leadership development evaluation therefore moved from the question "Were the participants satisfied?" to "What decision was taken, and what result was achieved?" The program came to be measured by its business impact rather than by attendance. This is the lesson offered by corporate universities: when they are built around actual business problems, impact measurement becomes part of their design rather than a later burden.
Perhaps the best-known manifestation of this approach was GE's "Work-Out" program, launched in the late 1980s. It consisted of intensive sessions in which employees met to diagnose waste and bureaucracy and propose practical solutions, with the manager required to decide immediately - accept or reject - in front of everyone. The output went beyond a certificate of attendance to a decision made and a tangible improvement whose effect on cost, time, and work quality could be tracked. Measurement thus became a natural product of the design rather than an artificial addition.
Fifth: The Success Case Method and Predictive Analytics
Brinkerhoff developed the "Success Case Method", which compares those who successfully applied the training with those who did not, to identify what makes the difference in the work environment. This school later evolved towards "predictive learning analytics", which forecasts immediately after the program who is most likely and least likely to apply the learning, enabling intervention to reduce waste before it occurs (Phillips, 2016).
The essence of the method is that it does not settle for averages that conceal the truth. It deliberately targets the extreme cases: those who most successfully applied what they learned and those furthest from application, interviewing them to discover what produced success and what obstructed it. The answer is often found outside the training room - in manager support, opportunities to apply the learning, or incentives. Training rarely fails alone; the surrounding system fails with it. By comparing those who applied the learning with those who did not, the method approaches the logic of the control group discussed under the fourth pillar, giving measurement both documented, verifiable success stories and a candid diagnosis of obstacles (Brinkerhoff & Mooney, 2008).
Predictive analytics, in turn, moves measurement from a "post-mortem examination" to an "early warning". It captures initial signals, such as engagement and evaluation scores and the level of manager support, to estimate as soon as the program ends who is likely to apply the learning and who is at risk of regression. Support and reinforcement can then be directed to the latter group before the learning is lost. Measurement thus becomes an intervention tool that protects the return, not merely a record that documents the loss after it has occurred.
6. Towards a Comprehensive System for Measuring Training, Its Impact, and Its Return:
What pillars support a system that links training effort to its impact and return?
Building on these experiences, an integrated system that goes beyond superficial measurement can be formulated around six interconnected pillars,
summarized in the following table:
The details are as follows:
Pillar One - Backward Design:
Begin with the end in mind. The essence of this pillar is a reversal in the order of thinking. We do not begin with the question "What will we train?" but with "What change do we want to see in the workplace?" and then work backwards from it. This is not a recent innovation but an extension of a well-established educational tradition that has been repurposed for the training industry.
As early as 1949, Ralph Tyler articulated the foundational logic of this approach when he stated that educational objectives are the criteria by which materials are selected, content is organized, procedures are developed, and tests are prepared. Wiggins and McTighe later developed this logic within the "Understanding by Design" framework, known as "backward design".
It consists of three sequential stages:
- Identify the desired results.
- Determine the acceptable evidence that will demonstrate their achievement.
- Plan to learn experiences and instruction.
The decisive difference is that traditional design begins by selecting learning activities and then builds the evaluation around them. Backward design does exactly the reverse: it defines the desired outcomes before selecting training methods and measurement tools.
When this logic is applied to the levels of training evaluation, the ladder is turned upside down in the design of training content, even though it remains upward moving in measurement. We begin with the targeted organizational indicator (Level 4: What will change in the business?), derive from it the required workplace behaviors (Level 3), then identify the knowledge and skills that enable those behaviors (Level 2), and finally design the learning experience and its environment (Level 1). This is precisely what the Kirkpatrick school established in its principle that "the end is the beginning": the form of success is defined with stakeholders at the start of the initiative, not at its conclusion (Kirkpatrick & Kirkpatrick, 2016). With this sequence, measurement is born alongside analysis. The moment we define the target indicator; we have simultaneously defined what will be measured and how.
Action Mapping within This Principle:
The clearest practical embodiment of this principle in corporate training is "action mapping", developed by Cathy Moore in 2008. It combines performance consulting with backward design and focuses on real-world behaviors rather than test questions. Its starting point is explicit: define a business or organizational goal, then work backwards to identify the actions learners must perform to achieve it and the practice activities that support those actions, while removing all content that does not serve the required performance. Moore places her finger on the core measurement problem when she observes that our work is not seen as vital to the organization because most of what we do is not measured, and that placing a measurable organizational goal at the center demonstrates our value. The organizational goal is therefore not merely a destination; it is the evaluation criterion itself, enabling us to assess success and demonstrate the value of training.
The Practical Obstacle and How to Overcome It:
A recurring field challenge remains: most clients do not have a clear objective in the first place. The training designer must then move from implementer to performance consultant and elicit the objective through dialogue using a practical formula: identify an indicator the organization already measures and that our project can improve, then define the amount of improvement and the timeframe. This step alone moves us from training "requested for its own sake" to training "requested for its impact".
The conclusion of this pillar is that the superficiality of measurement described at the beginning of the article is not a defect in measurement tools, but in the timing of measurement thinking. When a program is designed from its intended result, the question "What changed?" becomes structurally answerable because the answer was written before the trainee entered the room.
Pillar Two - Distinguishing Effort Indicators from Impact Indicators:
This pillar rests on a fundamental distinction: not everything measured in training is evidence of its impact. Some indicators embellish reports without reflecting any genuine change in performance.
A class of indicators known in measurement literature as "vanity metrics" consists of figures that look attractive in reports but do not indicate genuine impact. Prominent examples include completion rates, training hours, and satisfaction scores. These are activity measures that capture commitment and attendance, not competence or knowledge transfer. The predicament can be summarized in one sentence: these indicators tell you that learning took place, not that it succeeded.
The size of the gap is documented. In a recent survey of learning and development professionals, 69% rated their data-analysis skills as "good" or "excellent", yet only 28.6% felt confident in demonstrating the business impact of their training, and only 12.8% reported tracking Return on Investment or cost savings.
The solution is not to eliminate effort indicators but to place them in their proper position. Indicators progress from activity indicators (Did they attend?), to learning indicators (Did they understand?), to behavioral indicators (Are they working differently?), and then to business-impact indicators (time to competence, error reduction, and revenue per employee). In management measurement, we also distinguish between two types: leading indicators that predict success by measuring early inputs, such as participation and assessment scores, and lagging indicators that reflect longer-term outcomes, such as productivity growth and customer satisfaction. Wisdom lies in using both. The more mature position is not to dismiss learning-value indicators as vanity metrics, but to treat them as leading indicators and rungs on a ladder to be tracked internally, not as achievements to be paraded before leadership.
Pillar Three - Measurement over Time (Before/During/After):
This pillar is grounded in a robust psychological fact: both learning and behavior erode if they are not maintained. Measurement conducted now a program ends captures only a passing peak that soon declines. We must therefore track progress over an extended period: establish a pre-program baseline, measure during the program, and follow up after 30 and 90 days.
As early as 1885, psychologist Hermann Ebbinghaus identified what became known as the "forgetting curve". Contemporary readings estimate that a person forgets, on average, approximately 50% of new information within one hour of learning it and approximately 70% within one day; the average loss may reach approximately 90% of new information during the first week unless it is reinforced. These figures are not obsolete: a 2015 study replicated the curve and obtained results like Ebbinghaus's original data (Murre & Dros, 2015). If this is the fate of abstract knowledge, what of a skill that requires practice and reinforcement?
More dangerous than forgetting information is behavioral regression. Behavior-change research has identified a recurring pattern known as "triangular relapse": the desired behavior rises temporarily during an intervention, declines after the intervention ends, and then returns towards the baseline. The conclusion from the habit literature is clear: many interventions successfully change behavior in the short term, but people commonly revert to their old routines and habits once the training intervention ends. This is precisely the behavioral explanation of the "scrap learning" discussed earlier. The group that attempts application and then regresses - approximately 65% in Brinkerhoff's findings - falls within this critical period during the first few weeks (Brinkerhoff & Mooney, 2008).
Why 30 and 90 Days Specifically?
These windows are not arbitrary. On the one hand, experts in measuring behavioral change recommend tracking it at multiple points because this separates genuine change from temporary enthusiasm. In practice, this can be collected through a structured behavioral evaluation by the direct manager at 30, 60, and 90 days after training. On the other hand, habit-formation research indicates that establishing a new behavior takes from several weeks to several months (Lally et al., 2010). The 30-90-day window therefore covers precisely the phase in which the fate of the behavior is decided: will it become an established habit, or fade through regression?
The Methodological Implication for Measurement:
It follows that Level 3 (behavioral change) cannot be measured by course completion or by passing a final test. It requires observable follow-up overtime and input from the direct manager on what the trainee does in the workplace. Measurement thereby shifts from a "snapshot" taken in the room to a "film" tracked in the field. The baseline reveals the starting point, measurement during delivery captures learning, and follow-up after 30 and 90 days reveals which learning survived erosion and which evaporated.
Pillar Four - Isolating the Effect of Training:
In complex systems, a result rarely arises from a single cause. Performance is influenced by leadership, systems, incentives, and market conditions as well as training. Accordingly, 68% of professionals reported that their greatest difficulty is understanding the specific effect of training because other factors within the company also affect results. Here lies the essential difference between the two schools: the Kirkpatrick model assumes that improvement resulted from the training program, whereas the Phillips model actively searches for other possible causes of the results.
Isolation Tools:
Isolation tools vary in precision and cost. The most prominent are control groups, trend-line analysis of performance data, predictive models, and estimates from participants, supervisors, and management of the proportion of impact attributable to training. The control group is the most precise: one group participates in the program while a comparable group does not, and the difference in their performance is attributed to the program. When properly designed, this is the most effective isolation method. Trend-line analysis, by contrast, projects the future value of an indicator as though the training had not occurred, then compares that projection with actual post-program data; the difference becomes an estimate of the learning effect. Two rules are required to ensure greater credibility:
First: combine more than one tool. Using several isolation methods together strengthens the attribution claim, provided the limits of the methodology and the confidence levels of the results are stated transparently. Second, estimate conservatively. Phillips requires estimates to be built on the most reliable and credible sources and costs and benefits to be calculated conservatively. Measurement thereby shifts from an easily refuted claim to an argument capable of withstanding scrutiny from financial management.
Pillar Five - Partnership with the Direct Manager:
In their seminal work on the "transfer of training" (1992), Mary Broad and John Newstrom established a matrix that became foundational in this field. It combines the time dimension (before, during, and after) with the role dimension (manager, trainer, and trainee) in a nine-cell matrix. One of its most important findings is that the direct manager is involved in two of the three most influential combinations affecting transfer, revealing the scale of the manager's role in converting learning into practical behavior. They therefore described the manager as the "manager of the transfer process" itself, not merely the person who sends an employee to the training room.
The danger lies in the critical period immediately after training. When employees return to work, they need sustained motivation, support from their supervisors, and reinforcement of the concepts. Without these, even the most engaged participants revert to their old habits. Quantitative studies confirm this: trainees who perceive strong support from their direct supervisors for participating in training and applying what they learned are more likely to initiate transfer and application.
The partnership can be activated by assigning the manager explicit responsibility across all three phases. Before the program, the manager meets the trainer to discuss the content, define training objectives, make time available for preparation, and encourage attendance. The manager's role continues and becomes decisive during and after the program. Responsibility for impact thereby moves from the shoulders of the trainee and trainer alone - where it is commonly and mistakenly placed - to where it belongs: the direct line of leadership, which can extend the leverage of training or cut it short.
Pillar Six - A Dual Language of Value (ROE and ROI):
This pillar begins with a practical premise that is often overlooked: not every stakeholder asks the same question. The operational sponsor asks, "Did this training achieve what we expected?" while the finance executive asks, "Was it worth what we spent on it?" The common error is to answer both questions with one measure. Results should instead be translated into two distinct languages, each used in its proper place.
First - Return on Expectations (ROE):
The Kirkpatrick school developed this concept to translate the value of training into the language of what stakeholders hope to achieve, not into the language of money alone. It is founded on the explicit principle that "the end is the beginning". At the outset of any training initiative, stakeholders jointly define what success will look like in observable or measurable terms - the intended profile of a program graduate. The school therefore regards Return on Expectations as the highest indicator of value.
The essence of Return on Expectations is that it begins clearly at Level 4. Rather than imposing a predetermined measure, stakeholders themselves define what they regard as success and agree on it in advance. This is a process of negotiation and clarification in which learning professionals ask enough questions to translate leadership's general expectations into observable, measurable results. The subtle difference between ROE and financial return is that ROI seeks to isolate the value of training alone, whereas ROE seeks to create the value that stakeholders care about, while acknowledging that training contributes only partially and that achieving the result requires an integrated system rather than a single program.
An important terminological caution is required: this "Return on Expectations" (ROE) must not be confused with the common financial meaning of the same abbreviation, "Return on Equity", which is measured as net income divided by shareholders' equity.
Second - Return on Investment (ROI):
Jack Phillips added a fifth level above Kirkpatrick's four levels, comparing business-impact results with the program's total fully loaded cost and revealing the net monetary benefit for every Saudi riyal spent. The calculation proceeds through six steps: identify improvement in impact indicators; isolate the portion attributable to the program; convert it into monetary value; calculate the full loaded cost; identify intangible benefits; and finally compare benefits with costs in the ROI ratio. The defining feature of the Phillips methodology is its insistence on isolating the effect of training from all other influencing factors through control groups, trend lines, and expert estimates. For example, if benefits amount to 90,000 against a cost of 50,000, ROI is 80%. The benefit-cost ratio is 1.8, meaning that each monetary unit invested returns an additional net 0.8 after the original cost has been recovered.
It is important to recognize that financial return is calculated only after verifying that impact occurred. This step depends on Level 4 impact indicators that improved because of the program. Converting an unproven result into a financial figure is therefore building on emptiness. The same applies to benefits that cannot be monetized: not every Level 4 indicator should be converted into money. Some remain intangible benefits either because they cannot be converted credibly without excessive cost or because their intrinsic value makes monetization unnecessary.
The system therefore concludes with a simple practical rule: match the measure to the question. Return on Expectations is the answer when the sponsor's concern is operational; Return on Investment is the answer when the question is explicitly financial. Neither can stand without impact that has been achieved and demonstrated at Level 4, rather than merely promised.
Sixth - From Theory to Practice: Kaizen's System for Measuring Training Efforts:
How can these pillars become a practical system applied in the field every day?
No matter how pillars sound, a theoretical system remains a deferred promise unless it is embodied in an institutional practice managed day by day. Kaizen offers an applied model worthy of examination: it has translated the six pillars into an integrated system for measuring "training effort", not momentary satisfaction alone. The system extends across the time axis - before, during, and after training - and draws evidence from every party, not from the trainee alone.
At the Centre: A Permanent Review System with Four Perspectives:
At the heart of the system is a "permanent system" that reviews training results and evaluates the training effort in every program. It does not rely on a single voice but brings together evidence from four parties: the trainee, the supervisor, the trainer, and the direct manager. This "four-perspective view" gives practical form to the fifth pillar's principle that the direct manager is a partner in creating impact, not merely the person who sends an employee to the training room (Broad & Newstrom, 1992). At the same time, it approaches the multi-source logic that strengthens the attribution claim under the fourth pillar. Bringing estimates together from four perspectives narrows the margin of bias and reveals what a single perspective may miss.
Three Measurement Stages:
Measurement tools are distributed across three successive stages. A pre-training stage establishes the baseline through "awareness-level measurement" and a pre-test assessment that explores what the trainee already knows. A during-training stage monitors delivery quality and the level of mastery. A post-training stage measures impact after field application has been practiced. This distribution is precisely what the third pillar advocates when it rejects reducing measurement to a "snapshot" taken in the room and calls instead for a "film" tracked in the field (Kirkpatrick & Kirkpatrick, 2016).
The Six Tools for Measuring Training Effort:
The system is structured around six integrated tools, progressing from program inputs to its ultimate impact:
- Awareness-Level Measurement: A pre-program tool that uses exploratory questions to identify what the trainee already knows about the program’s subject matter, thereby establishing the baseline against which subsequent progress is measured.
- Training Content Evaluation: The content is assessed against five criteria—comprehensiveness, clarity, sequencing, coherence, and relevance to field realities—to ensure that what is taught is suitable for practical application rather than memorization alone.
- Trainer Evaluation: Trainers are assessed through selection criteria, qualification forms, participation in “Training of Trainers” programs, and monitoring before and during delivery, as trainer quality is a prerequisite for the quality of the resulting impact.
- Assessment of Skill Mastery: This evaluates both the cognitive and practical dimensions of mastery, shifting the measurement question from “Did the trainee attend?” to “Did the trainee achieve mastery?”
- Evaluation of Training Achievement Quality: This is conducted through a tool that verifies the level of achievement by monitoring awareness, motivation, and acquired skills, supported by practical models and application-based tools such as projects, case studies, and the “Training Achievement Workbook.”
- Impact Measurement Tool: A post-program tool that measures actual changes following field application and answers the central question around which this entire article revolves: “What changed because of the training?”
This system covers the entire time continuum, draws on multiple sources of evidence, establishes an awareness baseline before claiming impact, distinguishes between cognitive and practical mastery, and does not overlook the quality of the content and the trainer as inputs that shape the outcome. In doing so, it provides a clear practical embodiment of the first, third, and fifth pillars, while reflecting the essence of the second pillar by focusing on “measuring effort” rather than merely relying on momentary satisfaction.
It also culminates its outputs in a dual language of value: “Return on Expectations” addresses the operational sponsor, while “Return on Investment” addresses the financial decision-maker.
Getting Started: What Should I Do Tomorrow Morning?
To avoid the paralysis of perfection, do not attempt to fix everything at once. Select one high-impact program and apply the following steps:
- Define the destination first. Meet with the decision-maker before designing the training and agree on one business indicator that you intend to improve, the desired degree of improvement, and the timeframe within which it should be achieved.
- Design the content backwards. Begin with the behavior required in the workplace, then identify the skill that enables that behavior, and finally develop the training content—not the other way around.
- Establish the baseline. Measure awareness and performance before the program so that you have a point of comparison against which subsequent change can be identified.
- Engage the direct manager. Assign the manager a documented task before and after the program and schedule behavioral assessments at 30 and 90 days.
- Distinguish effort indicators from impact indicators. Translate them into “Return on Expectations” for the sponsor and “Return on Investment” for the finance function—and do not monetize an impact that has not been proven.
Conclusion
What changes when training is measured by its impact on the organization rather than by impressions formed in the training room?
The distance between training that earns applause from participants and training that transforms an organization is the same as the distance between measurement that captures impressions and measurement that captures impact. An industry worth hundreds of billions should not reduce its success to a satisfaction survey. Moving towards a comprehensive system is not a technical luxury; it is a condition for the survival of the learning function and for maintaining its credibility before leadership that has begun—rightly—to ask what return is being generated from its expenditure.
References
- Allied Market Research (2025) ‘Corporate training market to reach $805.6 billion globally by 2035 at 7.0% CAGR’, PR Newswire. Available at: https://www.prnewswire.com/news-releases/corporate-training-market-to-reach-805-6-billion-globally-by-2035-at-7-0-cagr-allied-market-research-302584896.html (Accessed: 22 July 2026).
- Brinkerhoff, R. O. and Mooney, T. (2008) Courageous training: Bold actions for business results. Oakland, CA: Berrett-Koehler.
- Broad, M. L. and Newstrom, J. W. (1992) Transfer of training: Action-packed strategies to ensure high payoff from training investments. Reading, MA: Addison-Wesley.
- Chief Learning Officer (2011) ‘Scrap learning and manager engagement’. Available at: https://www.chieflearningofficer.com/2011/03/29/scrap-learning-and-manager-engagement/ (Accessed: 22 July 2026).
- Ebbinghaus, H. (1913) Memory: A contribution to experimental psychology. Translated by H. A. Ruger and C. E. Bussenius. New York: Teachers College, Columbia University. Original work published 1885.
- IMD (2025) Re-imagining Crotonville: Epicenter of GE’s leadership culture (A). Lausanne: IMD Business School. Available at: https://www.imd.org/research-knowledge/strategy/case-studies/re-imagining-crotonville-epicenter-of-ge-s-leadership-culture-a/ (Accessed: 22 July 2026).
- Kirkpatrick, J. D. and Kirkpatrick, W. K. (2016) Kirkpatrick’s four levels of training evaluation. Alexandria, VA: ATD Press.
- Kirkpatrick Partners (2024) Kirkpatrick foundational principles. Available at: https://www.kirkpatrickpartners.com/wp-content/uploads/2024/03/kirkpatrick-foundational-principles.pdf (Accessed: 22 July 2026).
- Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W. and Wardle, J. (2010) ‘How are habits formed: Modelling habit formation in the real world’, European Journal of Social Psychology, 40(6), pp. 998–1009. https://doi.org/10.1002/ejsp.674
- LinkedIn Learning (2022) 2022 workplace learning report. Available at: https://learning.linkedin.com/resources/workplace-learning-report-2022 (Accessed: 22 July 2026).
- LinkedIn Learning (2025) 2025 workplace learning report. Available at: https://learning.linkedin.com/resources/workplace-learning-report (Accessed: 22 July 2026).
- McCall, M. W., Lombardo, M. M. and Morrison, A. M. (1988) The lessons of experience: How successful executives develop on the job. Lexington, MA: Lexington Books.
- Moore, C. (2017) Map it: The hands-on guide to strategic training design. Virginia: Montesa Press.
- Murre, J. M. J. and Dros, J. (2015) ‘Replication and analysis of Ebbinghaus’ forgetting curve’, PLOS ONE, 10(7), e0120644. https://doi.org/10.1371/journal.pone.0120644
- Phillips, J. J. (2003) Return on investment in training and performance improvement programs. 2nd edn. Burlington, MA: Butterworth-Heinemann.
- Phillips, K. (2016) ‘How much is scrap learning costing your organization?’, Association for Talent Development. Available at: https://www.td.org/content/atd-blog/how-much-is-scrap-learning-costing-your-organization (Accessed: 22 July 2026).
- Rcademy (2026) ‘How to assess training effectiveness using Kirkpatrick’s model’. Available at: https://rcademy.com/how-to-assess-training-effectiveness-using-kirkpatricks-model/ (Accessed: 22 July 2026).
- SkyQuest Technology (2025) Corporate training market size, share, and growth analysis. Available at: https://www.skyquestt.com/report/corporate-training-market (Accessed: 22 July 2026).
- Training Industry (2026) ‘The Kirkpatrick model’. Available at: https://trainingindustry.com/wiki/measurement-and-analytics/the-kirkpatrick-model/ (Accessed: 22 July 2026).
- Tyler, R. W. (1949) Basic principles of curriculum and instruction. Chicago, IL: University of Chicago Press.
- Valamis (2026) ‘Kirkpatrick model: Four levels of learning evaluation’. Available at: https://www.valamis.com/hub/kirkpatrick-model (Accessed: 22 July 2026).
- Wiggins, G. and McTighe, J. (2005) Understanding by design. 2nd edn. Alexandria, VA: Association for Supervision and Curriculum Development.
- Al-Anazi, Muqbil bin Mohammed, Shamsi, Mohammed Anas, and Ghosh, Abhijit (2023) "The impact of training on employee performance: An applied study of the Government Printing Press Authority in the Kingdom of Saudi Arabia", International Journal for Research Publication and Studies, 4(39), pp. 318-362. Available at: https://www.ijrsp.com/volume/issue-39/13/ (Accessed: 22 July 2026).
- Za'eemi, Mourad (2016) "Return on investment in training", Journal of Human Sciences, Freres Mentouri University Constantine, pp. 57-78. Available at: https://revue.umc.edu.dz/h/article/view/2316 (Accessed: 22 July 2026).
- Kachroud, Iman, and Riyadh, Abdelkader (2020) "Evaluating the effectiveness of training programmes using the Kirkpatrick model from the perspective of trainees at the Tebessa Cement Corporation", Al-Bashaer Economic Journal, 6(1), pp. 760-778. Available at: https://search.emarefa.net/detail/BIM-1041937 (Accessed: 22 July 2026).