Evaluation protocol

Evaluate AI-generated Salesforce documentation against the source.

Readable prose is easy to demonstrate. A reliable explanation has to survive a review of the conditions, branches and components it describes.

7 min read
The short answer

Freeze a small, representative metadata sample and write the expected facts before evaluating the generated output. Review each factual claim against that sample, record omissions separately from incorrect statements, and measure the time spent checking and correcting the result.

Use the same sample and questions when comparing approaches. This is a proposed evaluation method, not a published benchmark or a claim about Aprity's measured accuracy.

A repeatable evaluation protocol

  1. Freeze the input

    Record the org, environment, retrieval date, component versions and accessible metadata types. Choose a few components you already understand, including a branching Flow, a validation rule and an interaction across components.

  2. Write expected facts first

    Have a knowledgeable reviewer list the trigger, conditions, fields read or written, branches and referenced components. Separate facts in the source from business context supplied by people.

  3. Review the generated explanation

    For every important statement, locate the evidence. Mark an incorrect statement as an error, an expected fact that was omitted as an omission, and a claim that cannot be assessed as unresolved. Keep those categories separate.

  4. Measure human work

    Record generation duration separately from review and correction time. Retain the output before correction. Note whether different reviewers agree and which parts repeatedly require manual investigation.

  5. Repeat after a controlled source change

    In an authorised test environment, change one known condition and regenerate the documentation. Check that the affected explanation reflects the new source and that unrelated facts remain correct. This checks documentation refresh, not production behaviour.

Turn a validation rule into testable questions

For the synthetic validation-rule example linked from our documentation resources, ask which stage activates the rule, how blank and non-positive amounts are handled, and what happens outside that stage. Verify the explanation against the published formula before treating the expected cases as correct.

A statement that a rule is configured is different from proof that a particular transaction was blocked. Runtime questions need execution evidence from an authorised environment, including the actor and test inputs. A citation or confidence score cannot replace that evidence.

Report counts with their denominators

If you report supported claims, state how many claims were reviewed, how the sample was chosen and how unresolved items were handled. Do not combine omissions, incorrect statements and unreviewed components into one unexplained accuracy score.

The downloadable CSV starts with NOT_REVIEWED states and empty result cells. Complete it with observations; blank cells do not mean zero errors. Include limitations when sharing the outcome, and rerun the same protocol after a material change to the generator or source.

Use it with your team

Free working files

Download the documentation evaluation worksheet (CSV)

No registration required. Work locally and keep your organisation’s details private.

References and further reading

Where aprity fits

Apply the same review to Aprity

Choose a scope your team knows during the evaluation. Inspect the generated pages and component references, record errors and unresolved questions, and decide whether the result is useful for your documentation process.

Salesforce documentation with aprity →

  • Bring known components to the evaluation
  • Verify facts before judging writing style
  • Measure review effort alongside coverage
Questions

Evaluate AI-generated Salesforce documentation against the source. — FAQ

Does a source citation prove that an explanation is correct?

No. The reviewer must check that the cited source supports the particular statement, including its conditions and exceptions.

Can I compare two tools using different orgs?

The result would mix tool differences with differences in inputs. Use the same authorised source sample, questions and review criteria, and report any unavoidable differences.

Keep reading

Related guides

See it on your own org, not on a slide.

A free 14-day evaluation documents your org from a computed dependency graph — read-only, no credit card, and you review the output before deciding anything.

Start free trial
Read-only Purged after each scan EU & US