From Feason

How Should We Evaluate AI That Answers Christian Questions?

Artificial intelligence can produce a confident theological answer in seconds. Confidence is not the same as care. A model can invent a quotation, attach a real person’s name to t…

FeasonAugust 22, 2026 · 5 min read
Faith & Reason

Artificial intelligence can produce a confident theological answer in seconds. Confidence is not the same as care. A model can invent a quotation, attach a real person’s name to the wrong idea, flatten several Christian traditions into one, or hide uncertainty behind polished prose.

That creates a practical question for anyone building with AI: how should we evaluate a system that answers questions about Christianity?

Feason is developing an internal Christian AI benchmark candidate to make that question measurable. The goal is not to decide which model is most Christian. It is to test whether an AI can find, cite, and fairly represent Christian sources while remaining honest about the limits of its evidence.

Test the answer against evidence

Ordinary AI evaluations often reward an answer for matching a preferred phrase or reaching a predetermined conclusion. That is too narrow for theological questions. A responsible answer must show where its claims came from and distinguish what a source says from what the model infers.

The Feason candidate therefore evaluates several dimensions separately:

  • Did the system retrieve relevant passages and source types?
  • Do its citations point back to the evidence it was given?
  • Does it preserve the meaning and limits of those sources?
  • Does it attribute a claim to the correct tradition?
  • Can it explain genuine disagreement without caricature or false consensus?
  • Does it avoid unsupported claims and fabricated quotations?
  • Does it say when the available corpus does not support a complete answer?

This is the same evidence-first discipline behind Feason Context and our guide to verifying theological AI citations. The model is not asked to become an authority. It is asked to become answerable to inspectable evidence.

Do not turn Christianity into a score

Some things should not be collapsed into a leaderboard.

This evaluation does not measure faith, holiness, salvation, spiritual maturity, pastoral wisdom, or whether a model is “Christian.” It cannot determine which church is true or settle disputed doctrine by averaging scores. A benchmark can examine the quality of an AI’s evidence use. It cannot examine a soul.

That boundary is essential. Without it, technical language can make a category mistake sound scientific. A high score for citation fidelity means that a model used its evidence well in the tested cases. It does not make the model a theologian, pastor, confessor, or spiritual guide.

Version one is also intentionally text-only. Images, sermons, audio, video, pastoral conversations, and real formation outcomes require different expert-reviewed evaluations. They should not be implied by a score built from written source retrieval.

Christian disagreement needs to remain visible

Christian sources come from communities with histories, languages, confessions, and real disagreements. A Catholic source should remain Catholic. An Eastern Orthodox source should not be silently rewritten in Protestant vocabulary. A Baptist confession should not be treated as the view of every Christian.

Fair representation does not require pretending that all positions agree. It requires naming agreement where it exists, explaining disagreement accurately, and giving readers enough provenance to inspect the sources themselves.

The candidate treats several errors as critical failures: fabricating a source or quotation, assigning a view to the wrong tradition, presenting disagreement as consensus, and treating missing corpus coverage as proof that a tradition has no teaching. These failures matter because they can mislead a reader even when the surrounding paragraph sounds reasonable.

Compare the same model with and without evidence

The internal development suite currently has two tracks.

The retrieval track contains 210 development questions. It tests whether Feason can retrieve acceptable source passages, expected source categories, and basic provenance fields. These questions are visible to repository contributors, so they are development checks rather than a contamination-resistant public test.

The grounding-impact track contains 16 open-ended cases. The same model receives the same evaluation prompt twice: once without Feason evidence and once with a frozen Feason evidence packet. The answers are presented in balanced order to a separately configured judge and can then be reviewed by people who score both answers independently.

Freezing the evidence matters. If two models receive different source packets, their scores are not a clean comparison. The same principle applies to prompts, model versions, settings, rubrics, and corpus versions. A credible result must record all of them.

We are not publishing model rankings from these development cases. Automated scores remain provisional until blinded human review and judge calibration are complete.

Authority must be earned

Feason can build the tools, but Feason cannot make a benchmark authoritative simply by naming it one.

Before a public benchmark release, the work needs a documented source methodology, a verified coverage and licensing inventory, per-question expert adjudication, and a separate held-out set with contamination controls. It also needs reviewers from relevant Christian traditions, clear conflict-of-interest rules, predefined agreement thresholds, an appeals and correction process, and independent replication.

Approvals must be attached to exact versions and hashes. Otherwise a reviewer may approve one set of cases while a later result quietly uses another. Public results should disclose the evaluated model, judge, prompts, settings, corpus version, evidence hashes, human-review state, and known limitations.

This is slower than publishing a leaderboard. It is also the difference between a marketing score and a standard that other people can inspect, challenge, and improve.

Why this is valuable

AI product teams can use the evaluation to see whether grounding actually improves an answer. Christian publishers and ministry-technology teams can use it to test citation and attribution failures before readers encounter them. Researchers can study retrieval over religious sources without pretending that one dataset represents all of Christianity. Model providers can identify where fluent answers outrun their evidence.

The wider value is trust. A reader should be able to ask not only “Does this sound right?” but also “What supports this claim, whose position is it, and what might be missing?”

That is the direction Feason is building toward: AI that shows its work, represents disagreement fairly, and knows when its evidence is incomplete.

The current work remains an internal candidate. Developers can explore the underlying evidence layer through the Feason Context quickstart or inspect its capabilities through the Feason MCP server.

Built with care.

Product notes, research, engineering, and company stories from Feason — where faith and reason meet careful technology.

More from the Feason blog →

More from Feason

August 14, 2026

Guided Scripture Meditation With Narration and Stillness

Practice guided Scripture meditation on Feason with biblical passages, spoken guidance, deliberate silence, and an original ambient score.

August 13, 2026

An Online Christian Theology Course Built for Judgment

Explore Feason’s online Christian theology course: twelve rigorous seminars with source-based exercises, AI-assisted assessment, and oral defense.

August 13, 2026

Christian AI Grounding: Evidence Before Eloquence

Explore Christian AI grounding with Feason Context, an evidence API for retrieving inspectable Scripture and historical sources before generation.