AI Competency Evaluation Tool
Evaluate your competencies with clear, useful insight
NVIDIA: Nemotron 3 Super
Balanced Nemotron for demanding everyday work
NEW
FREE
Your prompt will appear here…
Your beautifully formatted article will appear here once you generate.
No history yet
Your generations will appear here. Sign in to save them permanently.
Can this person do the job unsupervised? That is the only question a competence decision has to answer, and it is a different question from whether they know the material or attended the training.
Knowledge is testable on paper. Competence has to be judged against a standard, by somebody, on evidence.
Short answer: The AI Competency Evaluation Tool is a free tool that judges capability against a defined standard. You paste the standard and the evidence, set how strict the judgement should be, and it returns a criterion by criterion assessment with a clear position on each one.
What is AI Competency Evaluation Tool?
It evaluates a person against a standard rather than a piece of work. The subject is what somebody can demonstrably do, assessed criterion by criterion, with the strictness matched to the consequence of getting it wrong.
Compliance as a criterion
Evaluation Criteria includes Compliance, which is the right frame when a standard has to be met rather than approached.
Very Strict available
Four strictness settings plus a slider, and safety related competence deserves the top of that range.
Checklist and rubric output
Both suit observed competence, where the assessment is a set of judgements rather than a mark.
Readiness as an option
Readiness frames the output as a sign off decision rather than a general appraisal.
Improvements per criterion
List Improvements attaches what is missing to each criterion, which is what a development conversation needs.
Why Use AI Competency Evaluation Tool?
Because competence decisions get made on proxies, and every proxy has a failure mode.
| The proxy used | What it actually shows | What a criterion based judgement adds |
|---|---|---|
| Training attended | Presence in a room | What they can now do |
| Time in role | Duration, not capability | Evidence against each criterion |
| A supervisor's impression | One person's confidence | A stated standard applied consistently |
| A written test | Knowledge about the task | Whether the task can be performed |
How Does AI Competency Evaluation Tool Work?
The site puts every tool on one working surface, and evaluation is what sits behind this button.
- Prompt box. Whatever is being evaluated goes here, pasted or described. Bring the standard and the evidence together, since neither decides anything alone.
- Model selector. The engine list opens with MSB AI, Anthropic Claude AI, Google Gemini and others on the menu.
- Advanced options. Ten controls behind a collapsed panel, documented below.
- Generate. Standard, evidence and settings travel through the instruction layer for evaluation.
- Result card. The judgement lands with its word count shown underneath.
- Export row. DOC, TXT and HTML, and DOC matters because a competence record usually has to be kept.
- Activity history. Earlier evaluations remain below with copy, listen, reuse, download and open result, which lets a reassessment be measured against the first one.
Step-by-Step Guide
- Paste the competence standard exactly as it is written.
- Describe the evidence: tasks observed, work produced, questions answered.
- Set Evaluation Criteria to Compliance or Readiness.
- Set strictness by consequence rather than by preference.
- Ask for a position on each criterion rather than an overall verdict.
- Generate, then read the criteria where the evidence is thin.
- Gather more evidence rather than deciding on what you have.
An evaluation is only sound if you can confirm that:
- ✅ The standard is pasted, not paraphrased
- ✅ Every criterion has evidence against it
- ✅ The evidence is observed performance, not attendance
- ✅ Strictness matches the consequence of being wrong
- ✅ A qualified person reviews anything with a legal dimension
Best Use Cases
- Workplace sign off where somebody will work unsupervised afterwards
- Probation decisions that need to be defensible
- Promotion cases assessed against a role profile
- Reassessment after a period of development
- Consistency checks where several assessors judge the same standard
Not all evidence carries the same weight, and a competence file made of the wrong kind is thin however thick it looks.
| Evidence | What it establishes | Weight |
|---|---|---|
| Task performed unprompted, observed | Capability under normal conditions | Strong |
| Task performed on request | Capability when cued | Moderate |
| Correct answer to a question about it | Knowledge of the task | Weak on its own |
| Training attended | Exposure to the content | None |
Advanced Options Guide
Ten controls. The two strictness settings are the ones to think about, because they encode how serious a mistake would be.
| Option | What it controls | Setting for competence |
|---|---|---|
| Evaluation Criteria | Overall, Quality, Accuracy, Completeness, Strengths, Weaknesses, Readiness or Compliance | Compliance against a standard, Readiness for a sign off |
| Strictness | Lenient, Standard, Strict or Very Strict | Very Strict where somebody could be harmed |
| Output Format | Score + Feedback, Detailed Report, Checklist, Strengths / Improvements or Rubric | Rubric or Checklist, since this is a set of judgements |
| Feedback Style | Constructive, Direct, Detailed, Encouraging or Actionable | Direct for the decision, Actionable for the development notes |
| Give a Score | Adds an overall mark | Off. Competence is met or not met per criterion |
| List Strengths | Names what is already demonstrated | On |
| List Improvements | Names what is missing per criterion | On |
| Age-Appropriate Language | Adjusts wording | Off for a workplace record |
| Strictness Level | Slider from 1 to 100 | Around 80 for safety related work |
| Custom Instructions | Free text up to 1000 characters | The standard, the evidence gathered, and what a near miss should count as |
Example Inputs
Karim assesses new machine operators. He opens the AI Competency Evaluation Tool with the standard and the observation notes together.
Standard, as written in our sign off document:
1. Completes pre start checks in the stated order.
2. Identifies and reports a fault without prompting.
3. Uses the emergency stop correctly on request.
4. Records output and downtime accurately.
5. Explains what to do if a guard is missing.
Evidence from three observed shifts:
1. Done correctly all three shifts, though prompted once
on the second.
2. Reported a bearing noise unprompted on shift three.
3. Demonstrated correctly, twice.
4. Two entries missing on shift one, none missing after.
5. Answered correctly, mentioned isolation but not the
reporting step.
Evaluation Criteria = Compliance
Strictness = Very Strict
Output Format = Checklist
Feedback Style = Direct
Give a Score = Off
List Strengths = On
List Improvements = On
Age-Appropriate Language = Off
Strictness Level = 80
Custom Instructions = Give me met or not yet met on each
of the five, not an average. Treat a prompted pre start
check as not yet met. Say exactly what evidence is missing
for any criterion you cannot pass.
Example Outputs
Four criteria came back met and one not yet met, which is a considerably more useful result than a percentage. Criterion five failed on a specific and fixable omission: the reporting step was missing from an otherwise correct answer, and that is a five minute conversation rather than a retraining programme.
The instruction about prompting mattered. A prompted pre start check counted as not yet met on shift two, and because the following shift was unprompted the criterion still passed on the evidence overall, with the reasoning visible rather than buried in an average.
Pro tip Ask for met or not yet met per criterion and refuse an overall score. Averaging a safety standard is how somebody gets signed off while failing the one criterion that would have mattered, and the phrase not yet met also frames a reassessment rather than a verdict.
Caution Competence assessment is frequently regulated, and in many trades only a qualified assessor can sign somebody off. This tool can structure evidence and expose gaps in it, and it cannot carry the decision. Where safety, licensing or accreditation is involved, treat the output as preparation for a qualified human judgement and nothing more.
Note A standard has to exist before anybody can be judged against it. If yours is informal or missing, the AI Assessment Creator builds the criteria and the boundary descriptions first, and that is the step this tool depends on.
What works well
- Judges per criterion rather than producing an average
- Names precisely what evidence is missing
- Strictness can be matched to real consequence
- Free to use, with no account needed
What to watch for
- It cannot sign anybody off, and in regulated trades only an assessor can
- Evidence quality decides everything and comes from you
- A paraphrased standard produces a judgement against the paraphrase
- Attendance and time served are not evidence of competence
AIToolsay is a free platform with a large library of AI tools, and this evaluator sits in the exam and assessment group. Each tool covers a single job, arrives with its own controls, and runs on prompt engineering written for it, which is why this judges a person against a standard while an assignment tool judges a document. Nothing installs and no account is needed. The engine menu spans MSB AI, OpenAI ChatGPT, Google Gemini, Anthropic Claude AI, DeepSeek and more, and a second reading is worth having before a borderline sign off. The full library sits behind the AIToolsay homepage.
Frequently Asked Questions
Is the AI Competency Evaluation Tool free?
Yes. There is no account step and nothing metered.
Can it sign somebody off as competent?
No. It structures the evidence and identifies gaps. The decision belongs to a person, and in regulated work to a qualified assessor.
Why avoid an overall score?
Because averaging lets somebody pass while failing a criterion that matters on its own. Met or not yet met per criterion is the honest structure.
What counts as evidence?
Observed performance, work produced, and answers given. Training attended and time in role are not evidence of capability.
How strict should I set it?
By consequence. Very Strict where a mistake could hurt somebody, Standard for a developmental review.
Can it help with consistency between assessors?
Yes, and that is a good use. Running the same evidence against the same standard gives a reference point when two assessors disagree.
Competence is a judgement about what somebody can do, made against a written standard, on evidence somebody gathered. The tool can make that structure explicit, which is most of what makes such a decision defensible.
So open the AI Competency Evaluation Tool, paste the standard word for word, describe what you actually observed, and ask for a position on every criterion. Thanks for reading, and I hope the gaps are small ones. If it becomes part of how you prepare a sign off, the AIToolsay community is open to you, our social accounts post each new tool as it lands, push notifications reach you first, and the newsletter carries guides much like this one.
Let AI Speak.