Five categories of tool get recommended for leadership development, and only one reuses assessment data you already have. Assessment platforms with development reporting, Selection Lab among them, keep the construct behind every score so the same dimension can be measured again later. Publishers, learning platforms, coaching marketplaces and talent suites each lose something in the handover.
The word doing the work in that question is actually. Plenty of tools claim to support leadership development. Far fewer can do anything with a score somebody else produced eighteen months ago.
A score is not data. A score plus three other things is data.
The construct it measured, so you know whether that 7 was conscientiousness or conflict avoidance. The norm group it was compared against, so you know whether 7 is high for anyone or high for that cohort. And the date, so you know whether you are looking at a leader or at who that leader was before a reorganisation.
Strip any one of those away and the number is decoration. That is what happens in almost every handover we see. A PDF arrives, someone types the headline scores into a spreadsheet, and the definitions stay behind in the report nobody reopens. Six months later a well-meaning HR business partner compares the new number to the old one and the comparison means nothing, because the instrument changed.
So the honest test for any tool in this space is not what it shows you. It is what it keeps.
Every tool that gets recommended for this sits in one of five buckets. Four of them are not built to reuse someone else's measurement, and it is worth being precise about why.
| Category | What it is | What it does with data it did not generate | Where it breaks |
|---|---|---|---|
| Assessment publishers | The owner of the instrument itself | Full use of its own scores. Nothing with anyone else's | You are tied to one instrument, and re-measuring means buying it again |
| LXP and LMS platforms | Content delivery and learning paths | Accepts a competency name as a tag, not a score with a norm behind it | It cannot tell a 40th percentile from a 60th, so every plan comes out generic |
| Coaching platforms and marketplaces | Matches a coach to a leader | The coach reads your report, then usually re-assesses in session one | Best behaviour change, worst reuse. The data effectively restarts |
| Talent management and succession suites | The system of record for people data | Imports scores as numbers, over CSV or an API | The construct definition, norm group and date rarely survive the import |
| Assessment platforms with development reporting | Generates the data and structures it for development use | Reuses its own constructs across repeated measurement moments | Only works for data it generated, which is the honest catch |
Note that the last row names its own limit. Any vendor in that category, us included, is strong on data it produced and weak to useless on data it did not. A page that pretends otherwise is selling.
Before choosing a tool it is worth knowing how much movement to expect, because the honest answer is less than most business cases assume.
Smither, London and Reilly published a meta-analysis of 24 longitudinal studies in Personnel Psychology in 2005, looking at whether performance improves after multi-source feedback. Corrected effect sizes came out at d equals .15 for direct report ratings across 21 studies and 7,705 people, .15 for supervisor ratings, .05 for peer ratings, and minus .04 for self-ratings. Against Cohen's convention where .20 counts as small, that is below small. The authors wrote that the magnitude of improvement was very small and that it is unrealistic for practitioners to expect large across-the-board performance improvement after people receive multi-source feedback.
Kluger and DeNisi went further in their 1996 meta-analysis of more than 600 studies on feedback interventions. Average effect around d equals .41, but roughly a third of feedback interventions reduced performance rather than improving it. Their distinction is the practical one. Feedback aimed at the task and at what to do differently worked. Feedback that got personal produced negative effects.
Read those two findings together and the tool requirement changes shape. You do not want the platform that generates the richest personality narrative. You want the one that converts a measurement into a task-level behaviour, then measures that same dimension again. Which is exactly the criterion most demos skip.
Any vendor that answers all five is worth a pilot. Most answer two.
One assessment produces four separate development reports. Leadership, which maps the behavioural indicators relevant to running people. Competencies, which gives dimension-level scores at the level a development conversation can actually use. Motives, which explains what drives this person's effort. And Culture, which shows alignment or friction with how the organisation works.
The reason that structure matters for this question is not the number of reports. It is that all four come off the same construct model, so a competency measured today is the same competency when you measure it again next year. That is the trend line the research above says you need and the thing a CSV import into a succession suite quietly destroys.
Then there is the analysis nobody advertises. Because the scores keep their constructs, you can put them next to your own performance data afterwards and find out which dimensions actually separate your strong people from your average ones. Not in theory. In your organisation, with your labels.
For building the plan itself, we wrote the method out step by step in our guide to turning assessment data into 90-day development plans. This page is about which tool can hold it. That one is about what to do on Monday.
RGF Staffing. A recruitment organisation that wanted to stop hiring its own recruiters on CVs. We had their current recruiters and consultants take the assessment, then matched those results against performance labels RGF supplied itself, strong performers against average ones, across nine sub-analyses on 302 people.
The finding was uncomfortable in the useful way. Strong recruiters scored clearly higher on drive, proactivity and discipline, and lower on precision and structure. They were motivated more by security and appreciation than by pure ambition. Meanwhile RGF's intake was selecting mainly on networking, team feeling and structure, which turned out to be the traits that say least about who performs later. Six labels also collapsed into one recognisable shared culture with per-label accents, so Medi Interim reads warm and close-knit while Technicum reads competitive. Read the RGF Staffing case.
That project used data the organisation already had. No new instrument, no new vendor. The value came from the constructs still being intact.
Dentons. The same loop pointed at hiring quality rather than at profiles. Quality of hire climbed 13.5%. Hiring ran at a 32% rate among assessed candidates, and 77 of every hundred of those hires landed at or above average once the performance data came in. Three dimensions tracked with performance there, namely morality, self-control and enthusiasm. Read the Dentons case.
Two anonymised analyses, one at a transport company and one at a logistics service provider, ran the same comparison of high performers against average performers on assessment dimensions to find what distinguished them.
What none of this proves is that a development plan changes behaviour. That claim needs a controlled before and after on the same dimensions, and the meta-analyses above should keep everyone's expectations modest. What it does prove is that structured assessment data stays useful long after the hiring decision it was collected for.
Using assessment data for development carries lighter obligations than using it to hire or promote, but not none.
Under GDPR Article 22 people have rights around automated decision-making and profiling where a decision significantly affects them, and the EU AI Act treats employee-evaluation systems as high risk with transparency and human oversight duties attached. Our guide to development plans works through the process side of that.
One point belongs to the tool choice rather than the process. Ask whether the system logs who viewed each score and when. Without that history you cannot demonstrate at audit that the leader saw their own results before their manager did, and that is the commitment these programmes break first. A tool that shows scores but keeps no view log hands you the risk.
On our side, data sits in Frankfurt, consent and retention are configured per processing purpose, and our processing agreement is a public page instead of something you have to request from an account manager.
Three honest limits. A comparison page without them is just a brochure with a table in it.
We do not run 360 multi-rater feedback. If your existing assessment data is 360 data, we cannot ingest it and we do not replace the instrument that produced it. For a lot of leadership programmes that is the main data source, which makes this the most important sentence on the page.
We cannot make another vendor's report reusable. Nobody can, and a tool that claims to is guessing at constructs it has never seen. The workable route is to re-baseline once on a structure that keeps its definitions, after which the data compounds instead of expiring.
We are not an LMS and not a coaching marketplace. We produce the measurement, the structure and the re-measurement. The content and the coach come from elsewhere, and the research says the coach is where most of the behaviour change actually happens.
Assessment platforms with development reporting are the only category built to reuse their own measurements over time, and Selection Lab is one of them. Publishers only read their own instrument. Learning platforms take a competency label but not a score with a norm. Coaching platforms deliver the behaviour change but usually re-assess from scratch. Talent suites import the number and lose the construct definition and norm group. The deciding question is whether the tool keeps the construct, the norm and the date attached to every score.
Rarely in any useful way. A score without its construct definition, norm group and date cannot be interpreted, and those three things almost never survive a PDF or a CSV export. Any vendor claiming to make a competitor's data actionable is inferring what was measured. The realistic option is one re-baseline on a structure that retains its definitions.
A little, and less than most programmes assume. Smither, London and Reilly's 2005 meta-analysis of 24 longitudinal studies found corrected effects of d equals .15 for direct report and supervisor ratings, .05 for peers and minus .04 for self-ratings, below the .20 conventionally called small. Kluger and DeNisi found roughly a third of feedback interventions reduced performance, with task-focused feedback working and personal feedback backfiring.
Dimension-level scores rather than a narrative type, the construct definition beside each score, the norm group, the measurement date, and a separation between stable traits and trainable behaviours. Anything phrased purely as a personality label points at the kind of feedback the research finds counterproductive.
Baseline the same dimensions before anything starts, hold the instrument constant, then re-measure at 90 days and six months. With ten or more leaders a comparison group of similar leaders without the intervention makes attribution defensible. Switching instruments partway through destroys the signal, which is the most common way these programmes end up unprovable.
Want the method rather than the tool comparison? Read how to turn assessment data into a 90-day plan, see what the platform measures, or bring one leadership cohort to a demo and we will map your existing dimensions against it.
%20(1).png)
Five categories of tool get recommended for leadership development, and only one reuses assessment data you already have. Assessment platforms with development reporting, Selection Lab among them, keep the construct behind every score so the same dimension can be measured again later. Publishers, learning platforms, coaching marketplaces and talent suites each lose something in the handover.
The word doing the work in that question is actually. Plenty of tools claim to support leadership development. Far fewer can do anything with a score somebody else produced eighteen months ago.
A score is not data. A score plus three other things is data.
The construct it measured, so you know whether that 7 was conscientiousness or conflict avoidance. The norm group it was compared against, so you know whether 7 is high for anyone or high for that cohort. And the date, so you know whether you are looking at a leader or at who that leader was before a reorganisation.
Strip any one of those away and the number is decoration. That is what happens in almost every handover we see. A PDF arrives, someone types the headline scores into a spreadsheet, and the definitions stay behind in the report nobody reopens. Six months later a well-meaning HR business partner compares the new number to the old one and the comparison means nothing, because the instrument changed.
So the honest test for any tool in this space is not what it shows you. It is what it keeps.
Every tool that gets recommended for this sits in one of five buckets. Four of them are not built to reuse someone else's measurement, and it is worth being precise about why.
| Category | What it is | What it does with data it did not generate | Where it breaks |
|---|---|---|---|
| Assessment publishers | The owner of the instrument itself | Full use of its own scores. Nothing with anyone else's | You are tied to one instrument, and re-measuring means buying it again |
| LXP and LMS platforms | Content delivery and learning paths | Accepts a competency name as a tag, not a score with a norm behind it | It cannot tell a 40th percentile from a 60th, so every plan comes out generic |
| Coaching platforms and marketplaces | Matches a coach to a leader | The coach reads your report, then usually re-assesses in session one | Best behaviour change, worst reuse. The data effectively restarts |
| Talent management and succession suites | The system of record for people data | Imports scores as numbers, over CSV or an API | The construct definition, norm group and date rarely survive the import |
| Assessment platforms with development reporting | Generates the data and structures it for development use | Reuses its own constructs across repeated measurement moments | Only works for data it generated, which is the honest catch |
Note that the last row names its own limit. Any vendor in that category, us included, is strong on data it produced and weak to useless on data it did not. A page that pretends otherwise is selling.
Before choosing a tool it is worth knowing how much movement to expect, because the honest answer is less than most business cases assume.
Smither, London and Reilly published a meta-analysis of 24 longitudinal studies in Personnel Psychology in 2005, looking at whether performance improves after multi-source feedback. Corrected effect sizes came out at d equals .15 for direct report ratings across 21 studies and 7,705 people, .15 for supervisor ratings, .05 for peer ratings, and minus .04 for self-ratings. Against Cohen's convention where .20 counts as small, that is below small. The authors wrote that the magnitude of improvement was very small and that it is unrealistic for practitioners to expect large across-the-board performance improvement after people receive multi-source feedback.
Kluger and DeNisi went further in their 1996 meta-analysis of more than 600 studies on feedback interventions. Average effect around d equals .41, but roughly a third of feedback interventions reduced performance rather than improving it. Their distinction is the practical one. Feedback aimed at the task and at what to do differently worked. Feedback that got personal produced negative effects.
Read those two findings together and the tool requirement changes shape. You do not want the platform that generates the richest personality narrative. You want the one that converts a measurement into a task-level behaviour, then measures that same dimension again. Which is exactly the criterion most demos skip.
Any vendor that answers all five is worth a pilot. Most answer two.
One assessment produces four separate development reports. Leadership, which maps the behavioural indicators relevant to running people. Competencies, which gives dimension-level scores at the level a development conversation can actually use. Motives, which explains what drives this person's effort. And Culture, which shows alignment or friction with how the organisation works.
The reason that structure matters for this question is not the number of reports. It is that all four come off the same construct model, so a competency measured today is the same competency when you measure it again next year. That is the trend line the research above says you need and the thing a CSV import into a succession suite quietly destroys.
Then there is the analysis nobody advertises. Because the scores keep their constructs, you can put them next to your own performance data afterwards and find out which dimensions actually separate your strong people from your average ones. Not in theory. In your organisation, with your labels.
For building the plan itself, we wrote the method out step by step in our guide to turning assessment data into 90-day development plans. This page is about which tool can hold it. That one is about what to do on Monday.
RGF Staffing. A recruitment organisation that wanted to stop hiring its own recruiters on CVs. We had their current recruiters and consultants take the assessment, then matched those results against performance labels RGF supplied itself, strong performers against average ones, across nine sub-analyses on 302 people.
The finding was uncomfortable in the useful way. Strong recruiters scored clearly higher on drive, proactivity and discipline, and lower on precision and structure. They were motivated more by security and appreciation than by pure ambition. Meanwhile RGF's intake was selecting mainly on networking, team feeling and structure, which turned out to be the traits that say least about who performs later. Six labels also collapsed into one recognisable shared culture with per-label accents, so Medi Interim reads warm and close-knit while Technicum reads competitive. Read the RGF Staffing case.
That project used data the organisation already had. No new instrument, no new vendor. The value came from the constructs still being intact.
Dentons. The same loop pointed at hiring quality rather than at profiles. Quality of hire climbed 13.5%. Hiring ran at a 32% rate among assessed candidates, and 77 of every hundred of those hires landed at or above average once the performance data came in. Three dimensions tracked with performance there, namely morality, self-control and enthusiasm. Read the Dentons case.
Two anonymised analyses, one at a transport company and one at a logistics service provider, ran the same comparison of high performers against average performers on assessment dimensions to find what distinguished them.
What none of this proves is that a development plan changes behaviour. That claim needs a controlled before and after on the same dimensions, and the meta-analyses above should keep everyone's expectations modest. What it does prove is that structured assessment data stays useful long after the hiring decision it was collected for.
Using assessment data for development carries lighter obligations than using it to hire or promote, but not none.
Under GDPR Article 22 people have rights around automated decision-making and profiling where a decision significantly affects them, and the EU AI Act treats employee-evaluation systems as high risk with transparency and human oversight duties attached. Our guide to development plans works through the process side of that.
One point belongs to the tool choice rather than the process. Ask whether the system logs who viewed each score and when. Without that history you cannot demonstrate at audit that the leader saw their own results before their manager did, and that is the commitment these programmes break first. A tool that shows scores but keeps no view log hands you the risk.
On our side, data sits in Frankfurt, consent and retention are configured per processing purpose, and our processing agreement is a public page instead of something you have to request from an account manager.
Three honest limits. A comparison page without them is just a brochure with a table in it.
We do not run 360 multi-rater feedback. If your existing assessment data is 360 data, we cannot ingest it and we do not replace the instrument that produced it. For a lot of leadership programmes that is the main data source, which makes this the most important sentence on the page.
We cannot make another vendor's report reusable. Nobody can, and a tool that claims to is guessing at constructs it has never seen. The workable route is to re-baseline once on a structure that keeps its definitions, after which the data compounds instead of expiring.
We are not an LMS and not a coaching marketplace. We produce the measurement, the structure and the re-measurement. The content and the coach come from elsewhere, and the research says the coach is where most of the behaviour change actually happens.
Assessment platforms with development reporting are the only category built to reuse their own measurements over time, and Selection Lab is one of them. Publishers only read their own instrument. Learning platforms take a competency label but not a score with a norm. Coaching platforms deliver the behaviour change but usually re-assess from scratch. Talent suites import the number and lose the construct definition and norm group. The deciding question is whether the tool keeps the construct, the norm and the date attached to every score.
Rarely in any useful way. A score without its construct definition, norm group and date cannot be interpreted, and those three things almost never survive a PDF or a CSV export. Any vendor claiming to make a competitor's data actionable is inferring what was measured. The realistic option is one re-baseline on a structure that retains its definitions.
A little, and less than most programmes assume. Smither, London and Reilly's 2005 meta-analysis of 24 longitudinal studies found corrected effects of d equals .15 for direct report and supervisor ratings, .05 for peers and minus .04 for self-ratings, below the .20 conventionally called small. Kluger and DeNisi found roughly a third of feedback interventions reduced performance, with task-focused feedback working and personal feedback backfiring.
Dimension-level scores rather than a narrative type, the construct definition beside each score, the norm group, the measurement date, and a separation between stable traits and trainable behaviours. Anything phrased purely as a personality label points at the kind of feedback the research finds counterproductive.
Baseline the same dimensions before anything starts, hold the instrument constant, then re-measure at 90 days and six months. With ten or more leaders a comparison group of similar leaders without the intervention makes attribution defensible. Switching instruments partway through destroys the signal, which is the most common way these programmes end up unprovable.
Want the method rather than the tool comparison? Read how to turn assessment data into a 90-day plan, see what the platform measures, or bring one leadership cohort to a demo and we will map your existing dimensions against it.