Some people eat, sleep and chew gum, I do genealogy and write...

Thursday, April 9, 2026

The Main Challenges of Full-text Search Part One

 

Three of the major online family tree/data base websites have implemented AI based full-text search and to some degree, handwriting recognition in the last three or so years. FamilySearch.org's offering is called "Full Text Search" and includes handwriting recognition. The Full Text Search is available for free to all users. MyHeritage.com introduced a similar program called Scribe AI. Ancestry.com's contribution is confined to OCR and lacks handwriting recognition. All the efforts of the genealogy programs are behind the ability of Google Gemini with NotebookLM and some of the other AI websites. Of course none of the genealogy programs have the resources of Google or OpenAI and the others. 

As far as the genealogical community is concerned, handwriting recognition, document translation, and full-text search are revolutionary in changing the way we do research. I can put hundreds of documents into NotebookLM and then have a discussion with Gemini about the contents of all the documents at onece and the conversation quicky evolves into an examination of what records need to be found to resolve serious historical issues. But working with full-text search opens a whole new series of challenges. 

The first and major challenge is the issue of search terms. In the old way of searching using Google, for example, you can formulate search terms based on categories. For example, if I want to find some probate records, I search for probate records in the state or country where my ancestors lived. I didn't expect Google to find the actual records of my ancestors. That was left to me to search through the records individually for the people I was interested in finding. Of course, I could go back to the time when I was forced to use microfilm and read through the entire collection of records to find one name. The time and effort spent in searching record collections severely limited the amount of information I was able to gather in any reasonable time period. In addition, of course, I was faced with the issue of deciphering all of the handwritten records. The results of these conditions were that very little research was produced in any particular time spent. 

 Full-text search opens up the ability to review hundreds, perhaps thousands of records in a matter of minutes, but the catch is the fact that you need to guess the search terms to use to open up the information. Let's say your ancestor's name was John. Here is a list, generated by Google Gemini, of the variations of the name John. 

The name John is one of the most enduring and prolific names in history. It originates from the Hebrew name Yohanan (יוֹחָנָן), meaning "Yahweh is gracious."

Because of its biblical significance—associated primarily with John the Baptist and John the Apostle—the name spread through the Byzantine Empire, into the Latin-speaking West, and eventually across the globe through various linguistic shifts.


Linguistic Evolution

The path of the name from its ancient roots to modern English follows a clear phonological chain:

  1. Hebrew: Yohanan

  2. Greek: Ioannes (Ἰωάννης)

  3. Latin: Iohannes / Johannes

  4. Old French: Jehan

  5. Middle English: Jan / John


Variations by Region and Language

Western & Northern Europe

These variations reflect the Germanic and Romance linguistic families where the name has remained a top choice for centuries.

  • English: John, Jack, Johnny, Jon

  • French: Jean, Yannick (Breton)

  • Spanish: Juan, Juanito

  • Portuguese: João

  • Italian: Giovanni, Gianni, Nino, Vanni

  • German: Johannes, Hans, Jan, Johann

  • Dutch: Jan, Johan, Hanne, Hans

  • Scandinavian (Danish/Swedish/Norwegian): Johan, Jan, Jens, Hans

The British Isles (Celtic Variations)

The Celtic languages developed distinct versions of the name that have since become globally popular in their own right.

  • Irish: Seán, Shane, Shaun, Shawn

  • Scottish Gaelic: Iain, Ian

  • Welsh: Evan, Ioan, Ieuan, Siôn

  • Manx: Juan

Eastern Europe & Eurasia

In Slavic and Orthodox traditions, the name often maintains the "I" or "Y" sound from the original Greek Ioannes.

  • Russian: Ivan, Vanya

  • Polish: Jan, Janusz

  • Czech/Slovak: Jan, Ján, Janko

  • Hungarian: János, Jancsi

  • Romanian: Ion, Ioan, Ionuț, Nelu

  • Bulgarian/Serbian: Ivan, Jovan

  • Greek: Ioannis, Giannis, Yannis

Middle East & Africa

These versions often stem directly from the Hebrew original or the Islamic tradition.

  • Arabic: Yahya (يحيا), Yuhanna (يوهنا)

  • Hebrew: Yohanan (modern: Yochanan)

  • Amharic (Ethiopia): Yohannes

  • Turkish: Yahya

Asia & Pacific

In these regions, the name is often adopted through religious conversion or phonological adaptation of Western names.

  • Chinese: Yuēhàn (約翰)

  • Japanese: Yohane (ヨハネ - Biblical), Jon (ジョン)

  • Korean: Yohan (요한)

  • Hawaiian: Keoni


Diminutives and Medieval Short Forms

Historically, many surnames were created from pet names or shortened versions of John.

  • Hank: Derived from the Dutch Hanne.

  • Jan: Common in Northern Europe; used as a root for many surnames.

  • Jenkin: A medieval English diminutive ("Little John").

  • Hick/Hitch: Obsolete medieval English rhyming nicknames for John.


Summary Table of Major Forms

LanguagePrimary FormCommon Diminutive
EnglishJohnJack
SpanishJuanJuanito
RussianIvanVanya
GermanJohannesHans
ItalianGiovanniGianni
IrishSeánShane
ScottishIanIain
FinnishJukkaJani
 Which one of all of these terms was the one used by your ancestor named John? Did he use the name John at all, or did he use some other name, such as Bubba or Kid or J.T.? So when you are faced with a search field such as this one from FamilySearch, what are you going to use for the search terms?


If you assume that the person's name was John, what are your chances of finding him if he went by one of the other names?  For example, my great-grandfather's official name was Henry Martin Tanner, but when he signed legal documents, such as deeds, he always used Henry M. Tanner. Full-text searches are rather literal, and if I search for Henry Tanner. I will possibly not find Henry M Tanner.  I can use all sorts of Boolean algebraic terms, but I will still face the same problems of determining the search terms I need to use to find any specific piece of information I am searching for. Another example: one of my relatives is named Joseph Christiansen. His grave marker says Joe Christiansen. He apparently did not like to be called Joseph. How am I supposed to know this?

The people programming full-text search could add the variations for all the names and all the places and practically everything else into their program. They might even implement artificial AI to recognize all of the variations. What happens in that circumstance is that the number of documents discovered by AI can run into the millions.

Do I have a solution for this? No, but I have a methodology I use to attempt to narrow down the number of possible variations. This primarily includes carefully reviewing the documents that I do find to discover the possible limited variations of the name used by the person I am searching for. This whole process also applies to place names and, to some extent, to dates, particularly when you think about calendar changes.

This is part one of this particular series of articles, and hopefully you will stay with me and read the rest of the series as it comes out over the next few weeks.

Friday, April 3, 2026

Another cautionary tale: Lawyers fined for AI-generated Legal Documents

 

At first I thought the story of an attorney being sanctioned by the court for submitting a legal brief with AI fabrications had to be an urban legend, but recently I began to learn about the relatively large number of cases with the same theme. 

The case of Mata v. Avianca, Inc. Case 1:22-cv-01461-PKC   Document 54, United States District Court, Southern District of New York shows that the cases of the AI-generated briefs are not urban legends. The Court held in part:

In researching and drafting court submissions, good lawyers appropriately obtain assistance from junior lawyers, law students, contract lawyers, legal encyclopedias and databases such as Westlaw and LexisNexis.  Technological advances are commonplace and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance.  But existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.  Rule 11, Fed. R. Civ. P.  Peter LoDuca, Steven A. Schwartz and the law firm of Levidow, Levidow & Oberman P.C. (the “Levidow Firm”) (collectively, “Respondents”) abandoned their responsibilities when they submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT, then continued to stand by the fake opinions after judicial orders called their existence into question.

In my extensive court experience it was not uncommon to find an attorney citing a case that either did not apply to the issues before the court or did not exist at all and this was long before the existence of generative AI. According to the latest data from researchers tracking this trend (notably Damien Charlotin of HEC Paris), there have been over 1,200 documented instances worldwide, with approximately 800 of those occurring in U.S. courts. See Kaste, Martin. “Penalties Stack up as AI Spreads through the Legal System.” WGCU News | PBS & NPR for Southwest Florida, April 3, 2026. https://www.wgcu.org/2026-04-03/penalties-stack-up-as-ai-spreads-through-the-legal-system.

The irony of all these cases is that there are AI solutions that could be used to avoid the issues and still give substantial support to the judicial process. If the attorneys, or whomever, knew more about how to use AI and their own legal skills. all these cases could be avoided. However, I do have several suggested lessons that we should all learn. 

#1 AI is not a toy, it is a highly sophisticated tool and it takes a measurable amount of time and effort to learn how it can and should be used. 

#2 Read the fine print. This is actually Rule Ten of The Rules of Genealogy. See https://genealogysstar.blogspot.com/2025/10/another-new-rule-of-genealogy-for-2025.html Not only do you need to check your sources when doing genealogical research, you need to read the case law in the legal world. 

#3 String cites ask for trouble. If you don't understand this lesson, then you don't know a lot about the legal profession.

#4 Ultimately, you have to know how to ask AI questions and how to give commands and more than both of these, don't trust anything you haven't verified for yourself. 

That's probably enough suggested lessons for today. 

Thursday, April 2, 2026

Precision Inquiry: The real core of using AI for research

 

When the current wave of generative AI came out, the main draw seemed to be the Natural Language Interface (NLI). The NLI is not Voice Recognition. Voice Recognition is the "ears" (turning sound into text), while a Natural Language Interface (NLI) is the "brain" (understanding what that text actually means). But in a real sense, when you communicate with a computer, even with a NLI, you cannot assume that the computer "understands" what you are saying. You cannot have a functional "assistant" with only voice recognition, as it would just be a very fast typewriter that doesn't know how to follow orders.

In order to obtain reasonable research responses, it is important to understand how to ask AI questions and give AI directions. The accuracy and value of AI responses (chats) relies heavily on the structure and content of the directions it is given with which to respond. In short, you need to learn how to communicate with AI and realize that the level of accuracy of AI's response rests primarily on the instructions it is given. 

Fortunately, you can ask an AI chatbot how to ask questions and give directions. To get started, here are the two questions to ask AI:

How do I ask you questions?

How do I create a prompt?

Now, if you ask a question, before you click on the enter button, you might add another question:

Is the anyway to ask this question better?

If you draft a prompt or copy one from an online video or other source, ask the AI chatbot to suggest a better prompt. Many of the common AI programs have a way to assist you in clarifying your instructions. Google has an integrated program called Gems that appears on the Gemini window. You can ask Gemini for step-by-step instructions about how to write prompts using Gems. You can also ask Gemini to give you a prompt on a specific topic such as in-depth research or identifying photographs. 

Here is an example of prompt for identifying old photos that was written by Gemini from my earlier attempts.

Purpose and Role:

You are the "Photo ID Specialist," an elite AI persona possessing the combined expertise of a Master Photographer, a Social Historian, and a Board-Certified Genealogist. Your mission is to analyze historical imagery with forensic-level detail to help researchers date photos, identify subjects, and place ancestors within their correct historical and social context.

Step-by-Step Analysis Protocol:

When a user provides a photo or description, you must evaluate it through these four lenses:


Material & Process: Identify the medium (e.g., Daguerreotype, Tintype, Albumen Print, Cabinet Card, RPPC). Explain the chemical or physical indicators that led to this conclusion and the specific year range the process was prevalent.

Fashion & Grooming Forensics: Analyze clothing (collar shapes, sleeve widths, button styles, fabrics) and hairstyles. Use these as primary markers for dating within a 3–5 year window.

Studio & Social Context: Examine backdrops, props (e.g., "hidden mother" chairs, ornate pedestals), and studio stamps/imprints. Note the social class suggested by the subject's attire and "theatricality."

Provenance & Epigraphy: Analyze any handwritten notes, photographer marks, or postal stamps (for postcards) to narrow down the geographic location and specific timeframe.

Rules for Interaction:


Request the "Reverse": Always ask if an image of the back of the photo is available, as the mount or handwritten notes are often more "evidentiary" than the image itself.

The "Evidence First" Rule: Before giving a final date estimate, list the visual observations (the "evidence") that support your conclusion.

Genealogical Bridge: Always suggest how the identified date/location helps narrow down specific census records or vital records for the user.

Output Structure (Required Format):

To maintain a professional and scholarly tone, organize your findings as follows:


### Executive Summary: A 1-sentence "Best Guess" for the date and location.

### Physical Analysis: Description of the photographic process and physical condition.

### Fashion & Visual Markers: Detailed breakdown of clothing/style cues.

### Historical Context: What was happening in that region/era that influenced this photo?

### Genealogical Recommendations: Specific "Next Steps" for the family researcher.

### Clarifying Questions: Ask for specific missing details (e.g., "Are there any tax stamps on the back?").

Tone and Voice:


Meticulous & Scholarly: Use precise terminology (e.g., "Leg-o-mutton sleeves," "Foxing," "Emulsion").

Empathetic: Acknowledge the sentimental value of these "shadows of the past."

Yes, it is long and complicated but it works. 

Wednesday, April 1, 2026

A Family History of the Irish Famine


 https://www.findmypast.com/a-family-history-of


https://www.youtube.com/@AFamilyHistoryOf

A Family History of the Irish Famine, continues the "A Family History of..." series from Findmypast.com hosted by genealogist and research specialist, Jen Baldwin. Here is a summary of the Press Release. 

In anticipation of the April release of the 1926 Irish Census, a powerful four-part podcast series titled A Family History Of The Irish Famine was launched on all streaming platforms on Tuesday, March 31, 2026. Hosted by genealogist and research specialist Jen Baldwin, the series examines the Great Famine through the life of her own ancestor, Archibald McKenzie. She is joined by guest historian Fiona Fitzsimons from Trinity College Dublin, who provides expert insight into the social and political forces that influenced Archibald’s journey.

Archibald’s story begins with his respectable life in County Cork before tracing his descent into poverty and his confrontations with the law. The narrative eventually follows his desperate bid to escape Ireland through the lesser-known chain-migration pathway. Using a variety of archival sources including census records, parish registers, and court documents, the podcast reconstructs Archibald’s world to provide an intimate portrait of resilience and the devastating human cost of this historical period.

This production is part of the weekly podcast A Family History Of..., which explores significant moments in British and Irish history through the lens of real families. The program has previously featured a series on wartime women and will continue with several upcoming themed series. Future installments are set to cover the 1926 General Strike, the Dardanelles Campaign of Gallipoli, and the Battle of the Somme. Listeners can also find bonus episodes and detailed research material on the podcast website.

The series is produced by Findmypast.com, a British and Irish genealogy platform owned by the family-run publishing company DC Thomson. By partnering with institutions like the British Library and The National Archives, the platform provides access to over a thousand years of historical records, helping individuals discover the stories and context of their ancestors' lives.