← All posts

Responsible AI in Genealogy: How PenParse Measures Against the CRAIGEN Principles

Genealogy has a particular reason to be suspicious of artificial intelligence. In most fields a confident, fluent, wrong answer is an inconvenience. In family history it is a fabricated ancestor, and once a fabricated ancestor enters a public tree it gets copied, and the copies get cited, and within a few years the error is load-bearing for hundreds of other people's research.

The community has noticed. There are grassroots threads warning about AI tools that overclaim, and at least one well-funded AI genealogy company has already collapsed. Against that background, the Coalition for Responsible AI in Genealogy has published five principles for responsible AI use: accuracy, disclosure, privacy, education, and compliance. They are short, they are sensible, and they are endorsed within the professional genealogical community.

We are not a member of CRAIGEN and we are not endorsed by them. We have no affiliation with the coalition at all. But the principles are published openly, anyone can hold themselves to them, and a standard nobody is willing to be measured against is not worth much. So this post goes through all five and reports where PenParse stands, including the one where we currently fall short.

Accuracy

> "AI can generate false, biased, or incorrect content. Therefore, members of the genealogical community verify the accuracy of the information with other records and acknowledge credible sources of content generated by AI."

The uncomfortable part of this principle for any vendor is that it asks you to make verification possible, which means telling users where your tool is weak.

We publish our failure rates. On sixteen historical German documents with human-verified transcriptions, PenParse scores 42 to 47 percent named-entity accuracy and roughly 30 percent character error rate on Kurrent script. On 17th-century chancery hands we are close to useless. The full benchmark, including the documents we lose on, is in our Kurrent accuracy report, and the modern handwriting benchmark is published separately.

The number that matters most is one we would rather not have. Our own confidence display overstates itself roughly 44 to 48 percent of the time on Kurrent, against about 16 percent on modern handwriting. In other words, on the material genealogists most need help with, our own uncertainty estimate is itself unreliable, and we would rather say so than let someone discover it inside their family tree.

This is why every transcription is colour-coded per word rather than delivered as flat text. The point is not decoration. It is to hand you the specific words to go and check against another record, which is precisely what this principle asks for.

Try PenParse free — 3 pages, no signup required

Try it free

Disclosure

> "Acknowledging the use of AI enhances trust. Therefore, members of the genealogical community disclose, as context requires, when AI materially influences the creation or modification of content."

Two things to report here, one good and one not.

The good one: we tell people when a competitor is the better choice. For research that is heavily weighted toward Kurrent, we publicly recommend Transkribus over our own product, and we say plainly that their free tier — 50 pages a month against our 10 — is more generous than ours. We would rather someone use the right tool than churn off ours after a bad first week.

The one where we fall short: our exports do not currently carry any marker identifying the text as AI-generated. If you export a transcription to DOCX or plain text today, nothing in that file records that it came from an AI tool. Paste it into a family history document a year later and the provenance is gone. That is exactly the situation this principle exists to prevent, and it is our gap, not the user's.

We are adding an optional provenance line to exports. Until that ships, if you are sharing a PenParse transcription — in a tree, a forum post, a society journal, or a family document — please label it as an AI-assisted transcription. We would rather ask you to do our job for a few weeks than pretend the gap isn't there.

Privacy

> "AI usage can lead to unintended data exposure, putting private information at risk of being publicly disclosed. Therefore, members of the genealogical community take reasonable measures to safeguard private information when using AI."

Family documents are not neutral data. They contain living people's names, addresses, medical details, and the occasional thing a family would prefer stayed in a drawer.

Our position, stated in full in our privacy policy:

  • Your uploaded images are never used to train AI or machine learning models, by us or by any third party.
  • Images go to third-party AI providers solely to produce your transcription, and those providers are contractually prohibited from using API inputs for model training.
  • Anonymous uploads are permanently deleted after 24 hours. If you have an account, your images are retained until you delete them or delete your account, and you can delete any document and its data at any time.
  • We use no third-party tracking or advertising cookies.

The honest caveat is that we are not a zero-knowledge system. Your images do travel to a third-party model provider to be read, because that is how the transcription happens. If a document is sensitive enough that this is unacceptable, no cloud transcription tool is the right answer, and we would rather you knew that before uploading than after.

Education

> "The use of AI creates new opportunities and risks. Therefore, members of the genealogical community educate themselves about AI to maximize its benefits and minimize its risks to their work."

The most useful thing we can contribute here is measurement, because almost nobody in this category publishes any.

Our benchmark posts exist to let a researcher form a realistic expectation before trusting anything: what error rates look like on modern handwriting versus historical German Kurrent, why a model produces fluent invented text instead of admitting it cannot read a word, and why confidence scores are trustworthy in-distribution and misleading outside it. If you are starting from scratch, our guide to reading old cursive handwriting covers what to try before reaching for any tool at all.

One finding is worth repeating on its own, because it saves people money: in our testing, photographing documents flat, in good light, and at full resolution improves results more than switching vendors does. That is advice against our own commercial interest and it is still the single highest-return thing most people can do.

Compliance

> "Adherence to legal and contractual obligations is essential for the ethical use of AI. Therefore, members of the genealogical community comply with contracts, terms of service, intellectual property laws, and data privacy regulations when using or creating AI tools."

The place this shows up most concretely for us is the benchmark itself. Our Kurrent ground truth is expert human transcription published by SLUB Dresden under CC BY 4.0, plus the CC BY-SA "Greetings From!" postcard dataset. Both are properly licensed and attributed, and there is no AI-generated ground truth anywhere in the set — a benchmark scored against another model's output would only tell us that two models agree with each other.

Why publish this at all

A vendor writing a post about its own ethics is not automatically worth reading, and we are aware of how this genre usually goes.

The reason to publish is that these five principles are checkable. The accuracy claims are in a benchmark with the documents and licences named. The privacy claims are in a policy you can read. The disclosure gap is a specific missing feature we have now committed to in public, which means you can hold us to it. If any of it turns out to be wrong, we would genuinely rather hear about it — write to privacy@penparse.com and we will correct the record here.

The standard we would like to be judged by is simple: not whether an AI tool ever gets things wrong, because it will, but whether it tells you where and how often, and whether it makes the errors easy for you to find.

Frequently Asked Questions

What are the CRAIGEN principles?

CRAIGEN is the Coalition for Responsible AI in Genealogy. It has published five principles for responsible AI use in genealogical research: accuracy, disclosure, privacy, education, and compliance. The full text is available at craigen.org.

Is PenParse a member of CRAIGEN?

No. We are not a member and we are not endorsed by CRAIGEN. We have no affiliation with the coalition. The principles are published publicly, and we have chosen to hold ourselves to them and report the results in public.

Does PenParse use my documents to train AI models?

No. Your uploaded images are never used to train AI or machine learning models, by us or by any third party. Images are sent to third-party AI providers solely to produce your transcription, and those providers are contractually prohibited from using API inputs for model training.

How long does PenParse keep my documents?

If you have an account, your images are retained until you delete them or delete your account. Anonymous uploads are automatically and permanently deleted after 24 hours.

How accurate is AI handwriting transcription for genealogy?

It depends heavily on the script. On modern handwriting our measured character error rate is about 5 percent. On historical German Kurrent it is roughly 30 percent, with named-entity accuracy between 42 and 47 percent. We publish both figures rather than only the favourable one.

Should genealogists disclose that a transcription was AI-generated?

Yes. CRAIGEN's disclosure principle asks that researchers acknowledge when AI materially influenced content, as context requires. A transcription produced by an AI tool is exactly that, and it should be labelled as such when it is shared, published, or added to a family tree.

Ready to try it?

Upload a handwritten image and get clean, editable text in seconds.

Try PenParse free — 3 pages, no signup required