Skip to content
HRaizon Subscribe

How to Draft a Recommendation Letter With AI, Then Fix What the Model Gets Wrong

A practical guide for HR teams and managers writing recommendation letters with AI tools: what the research says about bias in letters, the legal rules on references, and a template that survives scrutiny.

Share X in f
Priya Ellison

A recommendation letter is one of the few hiring documents still written by a person about a person. That is also why AI drafting tools are attractive: managers ask HR for help, HR has a dozen to turn around, and a language model can produce a plausible letter in seconds. The problem is that the model reproduces the two things that already make letters unreliable — vague praise and gendered language — and adds a third: it will happily invent accomplishments.

This guide covers what letters actually do in a hiring decision, the legal rules on what you can say, where AI-generated letters go wrong, and a template that keeps the writer accountable for the content.

What a recommendation letter is worth

Letters carry less predictive weight than writers and readers assume. A meta-analysis of letters of recommendation in college and graduate admissions (Kuncel, Kochevar and Ones, 2014) found that the predictive validity of letters in their usual free-form format is no higher than other predictors, with only a small incremental contribution to outcomes such as degree attainment. The authors argue the format is the problem: unstructured letters let writers pick what to emphasise, so two letters about equally strong people can look very different.

The practical implication for HR is that a letter’s value comes from specific, comparable information — role, tenure, concrete results, how the person compares to peers in the same job — rather than adjectives. That is also the information an AI tool cannot supply on its own.

The research on bias in letters

Letters are not neutral. Several peer-reviewed studies find consistent differences in how men and women are described, independent of performance:

  • A study of letters for academic positions published in the Journal of Applied Psychology (Madera, Hebl and Martin, 2009) found women were described with more “communal” terms (helpful, kind, warm) and fewer “agentic” terms (confident, ambitious, independent) than men, and that communal language was negatively associated with hiring decisions.
  • An analysis of 1,224 recommendation letters for geoscience postdoctoral fellowships from 54 countries, published in Nature Geoscience (Dutt et al., 2016), found women were about half as likely as men to receive an “excellent” letter (“brilliant scientist”, “scientific leader”) rather than a merely “good” one (“solid”, “very productive”). The gap held regardless of the recommender’s gender or region.

These are academic samples, and the effect sizes vary. But the pattern — women get warmth, men get competence — is the same one that shows up in model output, discussed next.

What AI drafting tools get wrong

They reproduce gendered language

Researchers at UCLA tested ChatGPT and Alpaca by asking each to generate reference letters for otherwise identical candidates with male or female names. The resulting EMNLP 2023 paper (Wan et al., “Kelly is a Warm Person, Joseph is a Role Model”) found the models described female-named candidates with more communal, likeability language and male-named candidates with more leadership and agentic language, mirroring the human pattern above. The effect appeared at both the word level and the sentence level, and the authors note that a generated letter with a gendered tilt can propagate bias downstream when it is used in an actual decision.

Newer models may behave differently; the paper tested 2023 systems. But no current vendor publishes audited results for reference-letter generation specifically, so treat any drafting tool as unaudited for this use.

They fabricate

A language model asked to “write a strong recommendation for a senior engineer” will produce achievements. If you did not supply them, they are invented. A letter containing a project that never happened or a number nobody measured is a misrepresentation, and under the legal rules below, a misrepresentation in a reference is the thing that creates liability.

They leak data

Pasting an employee’s performance review, compensation or health information into a consumer chatbot sends that data to a third party under whatever retention terms the tool has. Medical information is the clearest problem: under the Americans with Disabilities Act regulations, information about an employee’s medical condition or history must be kept in separate, confidential files and used only for permitted purposes. It has no place in a reference letter or in the prompt that produces one. Use an enterprise deployment with a data-processing agreement, or give the tool only the facts you intend to publish in the letter.

Legal rules on references

The rules differ by jurisdiction, and this is general information rather than legal advice. Two principles recur across most of them.

Say only what you can back up

In the United States, a reference is protected from defamation claims in most states by a qualified privilege, provided it is made in good faith. California’s version, in Civil Code §47(c), covers “a communication concerning the job performance or qualifications of an applicant for employment, based upon credible evidence, made without malice, by a current or former employer”. Other states have similar statutes; Florida §768.095, for example, makes an employer immune from civil liability for information disclosed to a prospective employer unless it is shown “by clear and convincing evidence” that the information was knowingly false or violated the employee’s civil rights. The privilege is lost if the statement is knowingly false or malicious, so the practical test is whether every factual statement in the letter traces to something documented: a review, a project record, a performance metric.

The same standard applies in the UK. Employers there usually have no duty to give a reference at all, but if they give one it must be “fair and accurate” and the employer “must be able to back up the reference, such as by supplying examples of warning letters”. A former employee who can show a reference was misleading or inaccurate, and that they lost an offer as a result, can sue.

Omitting a serious known risk can also create liability

A positive letter can be actionable if it leaves out something that makes it misleading. In Randi W. v. Muroc Joint Unified School District (1997), the California Supreme Court held that school administrators who wrote glowing, unreserved recommendations for a vice principal — while knowing of prior sexual misconduct complaints — could be liable for negligent misrepresentation to a student he later assaulted. The court’s reasoning was narrow: the letters affirmatively recommended him without reservation, and the omitted facts presented a foreseeable risk of physical harm to third parties. It does not require employers to volunteer every negative, but it does mean a writer who knows of a serious safety-related issue should either address it or decline to write an unqualified endorsement.

The simplest policy answer, and the one many US employers have adopted, is a neutral-reference rule: HR confirms title, dates and sometimes final salary, and anything beyond that is a personal reference from a manager, written in their own name. If your organisation allows substantive letters, the sections below assume the writer is accountable for the content.

References and data-protection rights

Under UK data-protection law, a reference given in confidence for employment or education purposes is exempt from the subject’s right of access (Data Protection Act 2018, Schedule 2, paragraph 24). That exemption covers the reference itself, not the underlying personnel records it draws on, and it does not license inaccuracy. In the US there is no equivalent general right of access to references, but some states give employees access to their personnel files, which may include copies.

A structure that holds up

A letter that a hiring team can actually use, and that a writer can defend, has five parts. Use this as the template; feed it to an AI tool section by section if you want drafting help, but fill in the facts yourself.

1. Who you are and how you know the person. Your title, the period you worked together, the reporting relationship. One or two sentences. This is what tells the reader how much weight your assessment deserves.

2. What they did. Role, tenure, scope — team size, budget, customers, systems. Plain description, no adjectives.

3. Two or three specific results. A project, what the person’s part in it was, what changed as a result, with numbers where you have them. “Reduced ticket backlog from 400 to 90 over two quarters by redesigning triage” is checkable; “a results-oriented self-starter” is not. This section is where fabrication risk lives, so every item should come from a record you could produce.

4. A comparison. How the person ranks against others you have managed in similar roles, stated directly: “among the three strongest analysts I have managed in ten years”. Comparative statements are the part of a letter readers find most informative and the part generic templates leave out.

5. Your recommendation and contact details. State the recommendation plainly and offer to take a call.

What to leave out: anything about health, family, age, religion or other protected characteristics; anything you heard second-hand; anything you cannot document.

Using an AI tool responsibly

If you use a drafting tool, the following steps address the failure modes above:

  1. Write the facts first. Put sections 1 to 4 in bullet points before opening the tool. The tool’s job is sentence construction, not content.
  2. Give it only what will appear in the letter. No review documents, no salary history, no medical or leave information.
  3. Ask for neutral, specific language. Instruct the tool to avoid personality adjectives and to describe actions and outcomes. Then check the draft against the communal/agentic split: if a letter for one person reads “warm, supportive, a pleasure to work with” and a letter for another reads “decisive, a natural leader”, and the underlying facts are similar, rewrite.
  4. Strike anything you did not supply. Every achievement, number, or project name in the draft must map to a bullet you wrote. Delete the rest.
  5. Have a second reader check it for bias and accuracy. Where HR reviews letters, use a short checklist: facts documented, no protected characteristics, comparable language across letters for comparable people, no unreserved endorsement where a known serious issue exists.

A blank-prompt request (“write a recommendation for Sam, a great product manager”) fails every one of these steps. The output will be fluent, generic, possibly gendered, and partly invented.

Declining to write

If you cannot write a positive, accurate letter, say no. A lukewarm letter hurts the candidate, and a letter that overstates to avoid awkwardness exposes you. A short reply — “I don’t think I’m the right person to write this for you” — is standard, and candidates generally prefer it to a weak letter. Where company policy limits references to title and dates, say so and point the candidate to HR.

Summary

Letters predict less than people think, and the free-form format is most of the reason. The research on gender bias in human-written letters is consistent, and the first controlled study of AI-generated letters found the same pattern in the model output. The legal standard in most jurisdictions reduces to one test: can you back up every statement. An AI tool can help with sentence construction once the facts are written down. It cannot supply the facts, the comparison, or the accountability, and a letter is only worth as much as those three things.