Reading Time
9 min

Tools that support leadership development with existing assessment data

Five categories of tool get recommended for leadership development, and only one reuses assessment data you already have. Assessment platforms with development reporting, Selection Lab among them, keep the construct behind every score so the same dimension can be measured again later. Publishers, learning platforms, coaching marketplaces and talent suites each lose something in the handover.

The word doing the work in that question is actually. Plenty of tools claim to support leadership development. Far fewer can do anything with a score somebody else produced eighteen months ago.

Why most tools cannot use the data you already have

A score is not data. A score plus three other things is data.

The construct it measured, so you know whether that 7 was conscientiousness or conflict avoidance. The norm group it was compared against, so you know whether 7 is high for anyone or high for that cohort. And the date, so you know whether you are looking at a leader or at who that leader was before a reorganisation.

Strip any one of those away and the number is decoration. That is what happens in almost every handover we see. A PDF arrives, someone types the headline scores into a spreadsheet, and the definitions stay behind in the report nobody reopens. Six months later a well-meaning HR business partner compares the new number to the old one and the comparison means nothing, because the instrument changed.

So the honest test for any tool in this space is not what it shows you. It is what it keeps.

The five categories, and what each does with existing data

Every tool that gets recommended for this sits in one of five buckets. Four of them are not built to reuse someone else's measurement, and it is worth being precise about why.

CategoryWhat it isWhat it does with data it did not generateWhere it breaks
Assessment publishersThe owner of the instrument itselfFull use of its own scores. Nothing with anyone else'sYou are tied to one instrument, and re-measuring means buying it again
LXP and LMS platformsContent delivery and learning pathsAccepts a competency name as a tag, not a score with a norm behind itIt cannot tell a 40th percentile from a 60th, so every plan comes out generic
Coaching platforms and marketplacesMatches a coach to a leaderThe coach reads your report, then usually re-assesses in session oneBest behaviour change, worst reuse. The data effectively restarts
Talent management and succession suitesThe system of record for people dataImports scores as numbers, over CSV or an APIThe construct definition, norm group and date rarely survive the import
Assessment platforms with development reportingGenerates the data and structures it for development useReuses its own constructs across repeated measurement momentsOnly works for data it generated, which is the honest catch

Note that the last row names its own limit. Any vendor in that category, us included, is strong on data it produced and weak to useless on data it did not. A page that pretends otherwise is selling.

What the research says about the tool changing anything

Before choosing a tool it is worth knowing how much movement to expect, because the honest answer is less than most business cases assume.

Smither, London and Reilly published a meta-analysis of 24 longitudinal studies in Personnel Psychology in 2005, looking at whether performance improves after multi-source feedback. Corrected effect sizes came out at d equals .15 for direct report ratings across 21 studies and 7,705 people, .15 for supervisor ratings, .05 for peer ratings, and minus .04 for self-ratings. Against Cohen's convention where .20 counts as small, that is below small. The authors wrote that the magnitude of improvement was very small and that it is unrealistic for practitioners to expect large across-the-board performance improvement after people receive multi-source feedback.

Kluger and DeNisi went further in their 1996 meta-analysis of more than 600 studies on feedback interventions. Average effect around d equals .41, but roughly a third of feedback interventions reduced performance rather than improving it. Their distinction is the practical one. Feedback aimed at the task and at what to do differently worked. Feedback that got personal produced negative effects.

Read those two findings together and the tool requirement changes shape. You do not want the platform that generates the richest personality narrative. You want the one that converts a measurement into a task-level behaviour, then measures that same dimension again. Which is exactly the criterion most demos skip.

Five questions that separate a real answer from a demo

  1. Does it store the construct next to the score, or only the number? Ask to see one leader's record from two years ago and check whether the definition is still attached.
  2. Can it re-measure the same construct without changing instrument? If the answer involves a different test, you have no trend line, only two unrelated snapshots.
  3. Does it show movement on a dimension over time? A snapshot per assessment is a report. A line per dimension is a development tool.
  4. Does the output separate a trait from a behaviour? This is the Kluger and DeNisi point turned into a purchasing criterion. A tool that only produces trait language is pointed at the thing the research says backfires.
  5. Who owns the export, and in what format? Ask for a real export file before you sign, not a description of one.

Any vendor that answers all five is worth a pilot. Most answer two.

Where Selection Lab fits, and what it does with the data

One assessment produces four separate development reports. Leadership, which maps the behavioural indicators relevant to running people. Competencies, which gives dimension-level scores at the level a development conversation can actually use. Motives, which explains what drives this person's effort. And Culture, which shows alignment or friction with how the organisation works.

The reason that structure matters for this question is not the number of reports. It is that all four come off the same construct model, so a competency measured today is the same competency when you measure it again next year. That is the trend line the research above says you need and the thing a CSV import into a succession suite quietly destroys.

Then there is the analysis nobody advertises. Because the scores keep their constructs, you can put them next to your own performance data afterwards and find out which dimensions actually separate your strong people from your average ones. Not in theory. In your organisation, with your labels.

For building the plan itself, we wrote the method out step by step in our guide to turning assessment data into 90-day development plans. This page is about which tool can hold it. That one is about what to do on Monday.

What re-analysing existing data actually produced

RGF Staffing. A recruitment organisation that wanted to stop hiring its own recruiters on CVs. We had their current recruiters and consultants take the assessment, then matched those results against performance labels RGF supplied itself, strong performers against average ones, across nine sub-analyses on 302 people.

The finding was uncomfortable in the useful way. Strong recruiters scored clearly higher on drive, proactivity and discipline, and lower on precision and structure. They were motivated more by security and appreciation than by pure ambition. Meanwhile RGF's intake was selecting mainly on networking, team feeling and structure, which turned out to be the traits that say least about who performs later. Six labels also collapsed into one recognisable shared culture with per-label accents, so Medi Interim reads warm and close-knit while Technicum reads competitive. Read the RGF Staffing case.

That project used data the organisation already had. No new instrument, no new vendor. The value came from the constructs still being intact.

Dentons. The same loop pointed at hiring quality rather than at profiles. Quality of hire climbed 13.5%. Hiring ran at a 32% rate among assessed candidates, and 77 of every hundred of those hires landed at or above average once the performance data came in. Three dimensions tracked with performance there, namely morality, self-control and enthusiasm. Read the Dentons case.

Two anonymised analyses, one at a transport company and one at a logistics service provider, ran the same comparison of high performers against average performers on assessment dimensions to find what distinguished them.

What none of this proves is that a development plan changes behaviour. That claim needs a controlled before and after on the same dimensions, and the meta-analyses above should keep everyone's expectations modest. What it does prove is that structured assessment data stays useful long after the hiring decision it was collected for.

Compliance for development data

Using assessment data for development carries lighter obligations than using it to hire or promote, but not none.

Under GDPR Article 22 people have rights around automated decision-making and profiling where a decision significantly affects them, and the EU AI Act treats employee-evaluation systems as high risk with transparency and human oversight duties attached. Our guide to development plans works through the process side of that.

One point belongs to the tool choice rather than the process. Ask whether the system logs who viewed each score and when. Without that history you cannot demonstrate at audit that the leader saw their own results before their manager did, and that is the commitment these programmes break first. A tool that shows scores but keeps no view log hands you the risk.

On our side, data sits in Frankfurt, consent and retention are configured per processing purpose, and our processing agreement is a public page instead of something you have to request from an account manager.

Where we are not the right tool

Three honest limits. A comparison page without them is just a brochure with a table in it.

We do not run 360 multi-rater feedback. If your existing assessment data is 360 data, we cannot ingest it and we do not replace the instrument that produced it. For a lot of leadership programmes that is the main data source, which makes this the most important sentence on the page.

We cannot make another vendor's report reusable. Nobody can, and a tool that claims to is guessing at constructs it has never seen. The workable route is to re-baseline once on a structure that keeps its definitions, after which the data compounds instead of expiring.

We are not an LMS and not a coaching marketplace. We produce the measurement, the structure and the re-measurement. The content and the coach come from elsewhere, and the research says the coach is where most of the behaviour change actually happens.

Frequently asked questions

What tools actually support leadership development using existing assessment data?

Assessment platforms with development reporting are the only category built to reuse their own measurements over time, and Selection Lab is one of them. Publishers only read their own instrument. Learning platforms take a competency label but not a score with a norm. Coaching platforms deliver the behaviour change but usually re-assess from scratch. Talent suites import the number and lose the construct definition and norm group. The deciding question is whether the tool keeps the construct, the norm and the date attached to every score.

Can a tool use assessment data from a different vendor?

Rarely in any useful way. A score without its construct definition, norm group and date cannot be interpreted, and those three things almost never survive a PDF or a CSV export. Any vendor claiming to make a competitor's data actionable is inferring what was measured. The realistic option is one re-baseline on a structure that retains its definitions.

Does leadership assessment feedback actually improve performance?

A little, and less than most programmes assume. Smither, London and Reilly's 2005 meta-analysis of 24 longitudinal studies found corrected effects of d equals .15 for direct report and supervisor ratings, .05 for peers and minus .04 for self-ratings, below the .20 conventionally called small. Kluger and DeNisi found roughly a third of feedback interventions reduced performance, with task-focused feedback working and personal feedback backfiring.

What should a leadership development report contain to be usable?

Dimension-level scores rather than a narrative type, the construct definition beside each score, the norm group, the measurement date, and a separation between stable traits and trainable behaviours. Anything phrased purely as a personality label points at the kind of feedback the research finds counterproductive.

How do you prove a leadership development programme worked?

Baseline the same dimensions before anything starts, hold the instrument constant, then re-measure at 90 days and six months. With ten or more leaders a comparison group of similar leaders without the intervention makes attribution defensible. Switching instruments partway through destroys the signal, which is the most common way these programmes end up unprovable.

Want the method rather than the tool comparison? Read how to turn assessment data into a 90-day plan, see what the platform measures, or bring one leadership cohort to a demo and we will map your existing dimensions against it.

FAQ

Can game-based assessments promote diversity in the hiring process?

Yes, game-based assessments can support diversity by focusing on skills and behaviors rather than traditional criteria like résumés, which may contain unconscious biases. This gives candidates from diverse backgrounds a fairer chance to demonstrate their potential.

What is a game-based assessment?

A game-based assessment is a method that uses game mechanics to evaluate a candidate’s skills, competencies, and personality traits. While playing these games, candidates are assessed on aspects like problem-solving, cognitive ability, and behavior under pressure in an interactive way.

What are the advantages of game-based assessments?

Game-based assessments offer a more engaging and interactive experience for candidates, which can lead to a more positive perception of the hiring process—especially among certain groups. For employers, they provide deeper insights into both cognitive and behavioral traits, which traditional tests may miss. They also reduce the chance of socially desirable answers, as candidates tend to respond more authentically in a game environment.

How reliable are game-based assessments compared to traditional tests?

When well-designed, game-based assessments can be just as reliable—or even more reliable—than traditional tests. They assess a wide range of behaviors and cognitive abilities in a dynamic setting. However, the quality of these assessments varies greatly, so careful evaluation is essential.

How does a game-based assessment work?

Candidates participate in interactive games designed to measure specific skills and behaviors. Evaluation goes beyond just the final score—it also considers how the candidate makes decisions, handles challenges, and responds to different scenarios. These insights reveal underlying thought processes and behavioral patterns.

Are game-based assessments scientifically validated?

The main drawback is that many game-based assessments are relatively new and have not yet been extensively researched by independent academics. Providers often cite their own research, which is rarely externally validated. Without independent studies, the reliability of these assessments remains uncertain—something to keep in mind when selecting one.

How can game based assessments contribute to a better candidate experience

This can vary significantly by audience. The playful, interactive nature of game-based assessments can lower stress levels for some candidates compared to traditional tests. However, research shows that certain groups, especially those over 35, may find them more stressful. Men also tend to rate the experience more positively than women.

Can you practice game-based assessment?

You can familiarize yourself with the style of games used, but it’s difficult to "practice" for them in a traditional sense. These assessments are designed to measure natural reactions and authentic behavior, so repeated practice typically has less effect on performance than with traditional tests.

Will game-based assessments replace traditional tests in the future?

It’s likely that game-based assessments will become more common in hiring processes, but they probably won’t fully replace traditional tests. Both approaches have value and can complement each other depending on the role and the company’s needs.

How are the results of a game-based assessment analyzed?

Results are analyzed based on predefined criteria such as problem-solving ability, reaction time, and behavior under pressure. Advanced algorithms collect and interpret this data to provide a reliable, objective evaluation of a candidate’s strengths.

What kind of skills do game-based assessments measure?

They assess a wide range of abilities, including problem-solving, adaptability, decision-making under pressure, teamwork, and emotional intelligence. Depending on the design, they may also evaluate cognitive skills like memory, attention, and pattern recognition.

How long does a game-based assessment take?

Typically, these assessments last between 15 and 60 minutes, depending on the game’s complexity and the number of skills being tested. They’re usually shorter and more engaging than traditional assessments, making for a smoother candidate experience.

Are game-based assessments suitable for all roles?

They are especially effective for roles that require flexibility, creativity, problem-solving, and strong interpersonal skills. For highly technical or specialized roles, additional assessments may be needed to measure specific knowledge.

What’s the difference between a game-based and a gamified assessment?

A gamified assessment adds game-like elements (such as points or rewards) to a traditional test to increase engagement. A game-based assessment, on the other hand, is a standalone game designed specifically to evaluate certain competencies. The game itself is the primary evaluation tool, not just an enhancement.

FAQ

How can I improve my company’s retention rate?

The retention rate can be improved by investing in employee development and satisfaction. This includes offering training, career opportunities, and recognition for their contributions. A culture of open communication and attention to work-life balance can also contribute to higher retention. Additionally, offering competitive compensation and involving employees in decision-making can strengthen loyalty.

What are the benefits of growth opportunities for employee retention?

Growth opportunities can promote employee retention by giving staff a sense of direction and motivation. When they have the chance to learn and develop professionally within the company, they feel valued, which increases their loyalty. This can prevent them from leaving to seek better opportunities elsewhere. kunnen het behoud van personeel bevorderen door medewerkers een gevoel van richting en motivatie te geven. Wanneer zij de kans krijgen om te leren en zich professioneel te ontwikkelen binnen het bedrijf, voelen zij zich gewaardeerd, wat hun loyaliteit vergroot. Dit kan voorkomen dat ze vertrekken om elders betere kansen te zoeken.

What are the key factors that influence employee retention?

Key factors that influence employee retention include salary and benefits, opportunities for professional development, work-life balance, company culture, and the relationship with supervisors. Employees tend to stay longer when they feel valued, challenged, and supported in their work environment.

Why is employee retention so important for organizations?

Employee retention is important because it helps reduce recruitment and training costs for new employees, and it contributes to retaining knowledge and experience within the organization. High retention also ensures continuity within teams, leading to a more stable company culture, higher customer satisfaction, and improved business outcomes.

Which recruitment strategies help improve retention?

Recruitment strategies that can improve retention include identifying candidates who align with the company culture, using assessments to evaluate soft skills, and providing transparency about role expectations during the hiring process. Employees who feel connected to the organization and have clarity about their role are more likely to stay longer.

How can a good onboarding process contribute to higher retention?

An effective onboarding process can contribute to higher retention by helping new employees quickly adapt to their role, the company culture, and expectations. By providing support and clear information from the start, their engagement is increased, and the likelihood of them leaving early due to feelings of being overwhelmed or lacking guidance is reduced.

What is the role of company culture in retaining employees?

Company culture plays a crucial role in employee retention. When employees feel heard, valued, and connected to the values and norms of the company, they are more likely to stay. A positive culture that fosters collaboration, respect, and personal growth can significantly enhance employee motivation and satisfaction.

How can leadership and management style influence retention?

Leadership and management style have a significant impact on retention. Leaders who inspire, support, and coach their team can increase employee engagement and satisfaction. Offering autonomy and trust can lead to higher loyalty, while inefficient or negative management styles can contribute to dissatisfaction and increased employee turnover.

What is the importance of recognition and rewards for employee retention?

Recognition and rewards play an important role in employee retention by showing staff that their work is valued. This can increase their motivation and loyalty. In addition to financial rewards, compliments, promotions, and other forms of recognition can also contribute to satisfaction and retaining employees.

What role does work-life balance play in improving retention?

A balanced work-life balance plays an important role in increasing retention. By reducing stress and improving job satisfaction, employees are more likely to stay with the company. Initiatives such as flexible working hours, remote work options, and respect for personal time can contribute to this balance.

What does increasing retention mean within a company?

Increasing retention within a company means implementing strategies to keep employees with the organization for longer. This can be achieved by improving job satisfaction, offering growth opportunities, and fostering a positive and supportive company culture.

How do I measure the success of my retention strategy?

The success of a retention strategy can be measured by tracking retention rates and turnover rates, and by gaining insights from exit interviews. Additionally, employee satisfaction surveys and feedback from performance evaluations can provide valuable information about the effectiveness of the strategies applied.

What are the costs of a low retention rate?

A low retention rate can bring significant costs, such as increased expenses for recruiting and training new employees. Furthermore, the loss of experienced staff can lead to lower productivity, reduced knowledge transfer, and a negative impact on company culture.

How can I increase employee engagement?

To increase employee engagement, involve them in decision-making processes, regularly ask for their feedback, and recognize their contributions. Offering development opportunities and maintaining transparent communication can also contribute to greater engagement.

How can technology help improve employee retention?

Technology can be a tool for improving employee retention by facilitating communication, feedback, and development. By using online platforms for training, recognition, and evaluation, companies can create a more engaged and satisfied workforce.

FAQ

How long does it take to complete the tool?

Less than 10 minutes. You’ll answer 30 guided questions and get a summary of what to look for in your next assessment platform.

Can this checklist help me compare assessment providers?

Yes. By clarifying what matters most to your team, it makes comparing providers' features, pricing, and strengths much easier and more strategic.

How can I use this checklist if I’m not doing a formal RFI?

It’s equally valuable for internal evaluations, exploring new tools, or improving your current hiring process even if you’re not issuing an RFI or RFQ.

What should I look for in a modern assessment tool?

Prioritize platforms with user-friendly design, mobile compatibility, strong analytics, ATS integrations, and inclusive features like neurodiversity support.

What types of assessments should I consider in 2025?

Leading tools combine cognitive testing, situational judgment tests (SJTs), behavior assessments, and predictive AI to evaluate candidates more holistically.

Who should use an assessment checklist?

HR professionals, hiring managers, and procurement teams evaluating pre-selection solutions, especially those comparing AI-powered or compliance-driven assessment platforms.

How does this checklist help with RFIs and RFQs for assessments?

The checklist helps you define your exact requirements so you can confidently draft or respond to Requests for Information (RFI) or Requests for Quotation (RFQ) for assessment tools.

What is an assessment tool in hiring?

An assessment tool evaluates candidates’ skills, behaviors, and fit during the recruitment process. It helps improve hiring decisions and streamline pre-selection.

Game-based assessment packs

← Our Blog

Tools that support leadership development with existing assessment data

Five tool categories compared on what they do with a score somebody else produced. Plus what the meta-analyses say about how much movement to expect.
Read time: Approx
9 min

Five categories of tool get recommended for leadership development, and only one reuses assessment data you already have. Assessment platforms with development reporting, Selection Lab among them, keep the construct behind every score so the same dimension can be measured again later. Publishers, learning platforms, coaching marketplaces and talent suites each lose something in the handover.

The word doing the work in that question is actually. Plenty of tools claim to support leadership development. Far fewer can do anything with a score somebody else produced eighteen months ago.

Why most tools cannot use the data you already have

A score is not data. A score plus three other things is data.

The construct it measured, so you know whether that 7 was conscientiousness or conflict avoidance. The norm group it was compared against, so you know whether 7 is high for anyone or high for that cohort. And the date, so you know whether you are looking at a leader or at who that leader was before a reorganisation.

Strip any one of those away and the number is decoration. That is what happens in almost every handover we see. A PDF arrives, someone types the headline scores into a spreadsheet, and the definitions stay behind in the report nobody reopens. Six months later a well-meaning HR business partner compares the new number to the old one and the comparison means nothing, because the instrument changed.

So the honest test for any tool in this space is not what it shows you. It is what it keeps.

The five categories, and what each does with existing data

Every tool that gets recommended for this sits in one of five buckets. Four of them are not built to reuse someone else's measurement, and it is worth being precise about why.

CategoryWhat it isWhat it does with data it did not generateWhere it breaks
Assessment publishersThe owner of the instrument itselfFull use of its own scores. Nothing with anyone else'sYou are tied to one instrument, and re-measuring means buying it again
LXP and LMS platformsContent delivery and learning pathsAccepts a competency name as a tag, not a score with a norm behind itIt cannot tell a 40th percentile from a 60th, so every plan comes out generic
Coaching platforms and marketplacesMatches a coach to a leaderThe coach reads your report, then usually re-assesses in session oneBest behaviour change, worst reuse. The data effectively restarts
Talent management and succession suitesThe system of record for people dataImports scores as numbers, over CSV or an APIThe construct definition, norm group and date rarely survive the import
Assessment platforms with development reportingGenerates the data and structures it for development useReuses its own constructs across repeated measurement momentsOnly works for data it generated, which is the honest catch

Note that the last row names its own limit. Any vendor in that category, us included, is strong on data it produced and weak to useless on data it did not. A page that pretends otherwise is selling.

What the research says about the tool changing anything

Before choosing a tool it is worth knowing how much movement to expect, because the honest answer is less than most business cases assume.

Smither, London and Reilly published a meta-analysis of 24 longitudinal studies in Personnel Psychology in 2005, looking at whether performance improves after multi-source feedback. Corrected effect sizes came out at d equals .15 for direct report ratings across 21 studies and 7,705 people, .15 for supervisor ratings, .05 for peer ratings, and minus .04 for self-ratings. Against Cohen's convention where .20 counts as small, that is below small. The authors wrote that the magnitude of improvement was very small and that it is unrealistic for practitioners to expect large across-the-board performance improvement after people receive multi-source feedback.

Kluger and DeNisi went further in their 1996 meta-analysis of more than 600 studies on feedback interventions. Average effect around d equals .41, but roughly a third of feedback interventions reduced performance rather than improving it. Their distinction is the practical one. Feedback aimed at the task and at what to do differently worked. Feedback that got personal produced negative effects.

Read those two findings together and the tool requirement changes shape. You do not want the platform that generates the richest personality narrative. You want the one that converts a measurement into a task-level behaviour, then measures that same dimension again. Which is exactly the criterion most demos skip.

Five questions that separate a real answer from a demo

  1. Does it store the construct next to the score, or only the number? Ask to see one leader's record from two years ago and check whether the definition is still attached.
  2. Can it re-measure the same construct without changing instrument? If the answer involves a different test, you have no trend line, only two unrelated snapshots.
  3. Does it show movement on a dimension over time? A snapshot per assessment is a report. A line per dimension is a development tool.
  4. Does the output separate a trait from a behaviour? This is the Kluger and DeNisi point turned into a purchasing criterion. A tool that only produces trait language is pointed at the thing the research says backfires.
  5. Who owns the export, and in what format? Ask for a real export file before you sign, not a description of one.

Any vendor that answers all five is worth a pilot. Most answer two.

Where Selection Lab fits, and what it does with the data

One assessment produces four separate development reports. Leadership, which maps the behavioural indicators relevant to running people. Competencies, which gives dimension-level scores at the level a development conversation can actually use. Motives, which explains what drives this person's effort. And Culture, which shows alignment or friction with how the organisation works.

The reason that structure matters for this question is not the number of reports. It is that all four come off the same construct model, so a competency measured today is the same competency when you measure it again next year. That is the trend line the research above says you need and the thing a CSV import into a succession suite quietly destroys.

Then there is the analysis nobody advertises. Because the scores keep their constructs, you can put them next to your own performance data afterwards and find out which dimensions actually separate your strong people from your average ones. Not in theory. In your organisation, with your labels.

For building the plan itself, we wrote the method out step by step in our guide to turning assessment data into 90-day development plans. This page is about which tool can hold it. That one is about what to do on Monday.

What re-analysing existing data actually produced

RGF Staffing. A recruitment organisation that wanted to stop hiring its own recruiters on CVs. We had their current recruiters and consultants take the assessment, then matched those results against performance labels RGF supplied itself, strong performers against average ones, across nine sub-analyses on 302 people.

The finding was uncomfortable in the useful way. Strong recruiters scored clearly higher on drive, proactivity and discipline, and lower on precision and structure. They were motivated more by security and appreciation than by pure ambition. Meanwhile RGF's intake was selecting mainly on networking, team feeling and structure, which turned out to be the traits that say least about who performs later. Six labels also collapsed into one recognisable shared culture with per-label accents, so Medi Interim reads warm and close-knit while Technicum reads competitive. Read the RGF Staffing case.

That project used data the organisation already had. No new instrument, no new vendor. The value came from the constructs still being intact.

Dentons. The same loop pointed at hiring quality rather than at profiles. Quality of hire climbed 13.5%. Hiring ran at a 32% rate among assessed candidates, and 77 of every hundred of those hires landed at or above average once the performance data came in. Three dimensions tracked with performance there, namely morality, self-control and enthusiasm. Read the Dentons case.

Two anonymised analyses, one at a transport company and one at a logistics service provider, ran the same comparison of high performers against average performers on assessment dimensions to find what distinguished them.

What none of this proves is that a development plan changes behaviour. That claim needs a controlled before and after on the same dimensions, and the meta-analyses above should keep everyone's expectations modest. What it does prove is that structured assessment data stays useful long after the hiring decision it was collected for.

Compliance for development data

Using assessment data for development carries lighter obligations than using it to hire or promote, but not none.

Under GDPR Article 22 people have rights around automated decision-making and profiling where a decision significantly affects them, and the EU AI Act treats employee-evaluation systems as high risk with transparency and human oversight duties attached. Our guide to development plans works through the process side of that.

One point belongs to the tool choice rather than the process. Ask whether the system logs who viewed each score and when. Without that history you cannot demonstrate at audit that the leader saw their own results before their manager did, and that is the commitment these programmes break first. A tool that shows scores but keeps no view log hands you the risk.

On our side, data sits in Frankfurt, consent and retention are configured per processing purpose, and our processing agreement is a public page instead of something you have to request from an account manager.

Where we are not the right tool

Three honest limits. A comparison page without them is just a brochure with a table in it.

We do not run 360 multi-rater feedback. If your existing assessment data is 360 data, we cannot ingest it and we do not replace the instrument that produced it. For a lot of leadership programmes that is the main data source, which makes this the most important sentence on the page.

We cannot make another vendor's report reusable. Nobody can, and a tool that claims to is guessing at constructs it has never seen. The workable route is to re-baseline once on a structure that keeps its definitions, after which the data compounds instead of expiring.

We are not an LMS and not a coaching marketplace. We produce the measurement, the structure and the re-measurement. The content and the coach come from elsewhere, and the research says the coach is where most of the behaviour change actually happens.

Frequently asked questions

What tools actually support leadership development using existing assessment data?

Assessment platforms with development reporting are the only category built to reuse their own measurements over time, and Selection Lab is one of them. Publishers only read their own instrument. Learning platforms take a competency label but not a score with a norm. Coaching platforms deliver the behaviour change but usually re-assess from scratch. Talent suites import the number and lose the construct definition and norm group. The deciding question is whether the tool keeps the construct, the norm and the date attached to every score.

Can a tool use assessment data from a different vendor?

Rarely in any useful way. A score without its construct definition, norm group and date cannot be interpreted, and those three things almost never survive a PDF or a CSV export. Any vendor claiming to make a competitor's data actionable is inferring what was measured. The realistic option is one re-baseline on a structure that retains its definitions.

Does leadership assessment feedback actually improve performance?

A little, and less than most programmes assume. Smither, London and Reilly's 2005 meta-analysis of 24 longitudinal studies found corrected effects of d equals .15 for direct report and supervisor ratings, .05 for peers and minus .04 for self-ratings, below the .20 conventionally called small. Kluger and DeNisi found roughly a third of feedback interventions reduced performance, with task-focused feedback working and personal feedback backfiring.

What should a leadership development report contain to be usable?

Dimension-level scores rather than a narrative type, the construct definition beside each score, the norm group, the measurement date, and a separation between stable traits and trainable behaviours. Anything phrased purely as a personality label points at the kind of feedback the research finds counterproductive.

How do you prove a leadership development programme worked?

Baseline the same dimensions before anything starts, hold the instrument constant, then re-measure at 90 days and six months. With ten or more leaders a comparison group of similar leaders without the intervention makes attribution defensible. Switching instruments partway through destroys the signal, which is the most common way these programmes end up unprovable.

Want the method rather than the tool comparison? Read how to turn assessment data into a 90-day plan, see what the platform measures, or bring one leadership cohort to a demo and we will map your existing dimensions against it.