Can AI generate assessment exercises that genuinely measure leadership and behaviour?

Recently I asked Copilot to provide me with HMICFRS report key headings. With reference to what it knows of my work, and because AI is nothing if not relentlessly helpful, it offered to go a step further, asking, “shall I create the following: scenario-based items; behavioural anchors aligned to NFCC behaviours?”.

Straight-away I could see how compelling an offer like that could be for FRS personnel tasked with creating tasks and exercises to use as part of promotion processes. The ability to generate exercises almost instantly and cut costs dramatically while giving complete control over the tasks and outputs feels like an opportunity almost too good to be true.

I was curious and concerned. Can it really do what it says it can do?

In a nutshell, the answer is no.

You know when you ask AI if it can create an image of your dream garden, and it responds with ‘yes absolutely!’…and then it presents a weird half job, not really following instructions you did give, but using artistic licence with the things you definitely didn’t ask for? It’s a bit like that. Only, perhaps, to the untrained eye, less obvious.

From the outside, the exercise design can look okay. There is some structure, and the words sound sort of right. But there is no substance. The content doesn’t link to a potential range of human behaviour. It’s just mapped onto other key words. And actually, a candidate could perhaps complete this exercise in a superficial way, but the exercise scope doesn’t allow for indications on their preferences, motivations, values or leadership skills in practice. The candidates would only have enough content and context to produce, well, more poorly evidenced words. (And of course, that was within the closed model of my own co-pilot, drawing on related content; without this source material the output would likely be even more flimsy).

You might wonder how this differs from the type of exercises which are professionally developed, and mainly it’s that the words are prompts which are designed to elicit behaviour. You have to understand what behaviour these prompts may elicit, and what the presence of, or absence of behaviour within that context might mean. It wouldn’t be reliable to only create opportunity for that once, so there needs to be an inter-woven pattern of prompts and likely behaviours to build up a sort of performance map. That map can then lead to developmental insights, personalised and implementable. As I said, AI does words. It doesn’t currently understand human behaviour beyond a simplistic, stimuli-response sort of model. And people tend to be a bit more complex than that.

AI is useful for brainstorming ideas, mind mapping possibilities to spark your own creativity. But the next stage, the depth, the psychology of what people do, why and how this matches to what we need them to be doing, just isn’t there. It’s a façade, like the homes on a film set, a convincing front and when you walk through the front door there is nothing there.

I would caution anyone using AI for this sort of work to have it sense checked by someone with more experience. It’s easy to be misled by the promise of what it can deliver, but the risks when it comes to recruitment, promotion and selection can impact everything from wellbeing, team functioning, service delivery, resources, organisational objectives and beyond.