← What we produce

How guidelines are developed

Process

Every CogGuide guideline follows the same six steps, modelled on how NICE develops clinical guidelines. Only invited experts decide the recommendations.

What a guideline is

CogGuide produces evidence-based, expert-reviewed best practice guidelines for frontier AI developers on how their products should protect users' cognition, including memory, belief and judgement.

The term follows the Institute of Medicine's definition of clinical practice guidelines: recommendations informed by a systematic review of evidence and an assessment of benefits and harms.

The guidelines are not formal standards. Mature guidelines can later be put forward to a recognised standards body.

01

Scoping

CogGuide picks topics where there is evidence of harm and a clear AI behaviour developers can change. A short scope document then sets out:

  • the context and users covered
  • the AI behaviours in question
  • three to six review questions
  • what is out of scope
  1. Example review question

    When a user discloses a crime they witnessed or experienced, which AI questioning behaviours increase the risk of memory distortion?

02

Expert panel

Each topic has its own panel of 8 to 12 experts and an independent chair. Membership criteria are published in advance.

Members qualify through peer-reviewed research on the topic, or senior professional practice in it, such as police interviewing or forensic psychology. Panels mix disciplines, for example memory science, law and AI safety.

All members publish declarations of interest. A member with a financial tie to an affected developer can join the discussion but does not rate the recommendations it would affect.

03

Evidence review

The CogGuide research team runs a systematic review for each review question, using a protocol registered in advance on the Open Science Framework.

Reviews cover AI-specific studies and transferable human research, such as the evidence base on interviewing witnesses. The panel receives each review as evidence tables with a plain summary.

04

Recommendations

The panel discusses the evidence and drafts recommendations. Each recommendation states:

  1. The behaviour

    What the model should do, or avoid doing.

  2. The evidence

    The research the recommendation rests on.

  3. The certainty

    How strong that evidence is, rated with GRADE as high, moderate, low or very low.

  4. The test

    A check developers can run to see whether their system follows the recommendation.

Where agreement is unclear, members rate the wording anonymously from 1 to 9. A recommendation is agreed when at least 75 percent rate it 7 or above. Unresolved disagreements are recorded in the published guideline.

05

External review

The draft goes to independent expert reviewers who are not on the panel, with four weeks to comment.

Frontier developers are invited to comment on technical feasibility only, and have no say in the recommendations. The panel answers every comment in a published response table and revises the draft where needed.

06

Publication and updating

CogGuide publishes each guideline alongside the evidence reviews, panel membership, declarations of interest and the comment response table.

Each guideline is checked every two years, or sooner if major new evidence appears. Mature guidelines can later be put forward to a recognised body, for example as a BSI Publicly Available Specification.

Sources

  1. Institute of Medicine (2011). Clinical Practice Guidelines We Can Trust. Washington, DC: The National Academies Press.
  2. National Institute for Health and Care Excellence. Developing NICE guidelines: the manual (PMG20).
  3. GRADE Working Group. Grading of Recommendations, Assessment, Development and Evaluation.

Get involved

CogGuide welcomes researchers and practitioners who want to join a guideline panel or review a draft guideline.

Get in touch on LinkedIn