Some people eat, sleep and chew gum, I do genealogy and write...

Tuesday, March 10, 2020

Ancestry.com loads up with new collections


Likely in response to the recent setback in DNA testing, Ancestry.com has come out with a long list of new record collections. You can view all the new and updated collections in the Ancestry Card Catalog available in the Search menu at the top of most of the pages. Make sure you sort by date added.


Looking down through this long list of new collections, I see some significantly large collections including the following:


  • Ireland, Petty Session Court Registers, 1818-1926 Court, Land, Wills & Financial 23,437,916
  • Ireland, Prison Registers, 1790-1924 Court, Land, Wills & Financial 3,127,594
  • New York, New York, Index to Birth Certificates, 1866-1909 Birth, Marriage & Death 6,091,199
My work is always cut out for me with record hints. At this time, I have over 17,000 waiting to be processed. 

Monday, March 9, 2020

Reinventing Genealogy

An old man, identified as the legendary Flavio Amalfitano, seated at his desk using a compass and reading a book; tools, books, and instruments are arranged around the richly furnished room; a dog is at his feet. A model of a ship hangs above and in front of him. Jan van der Straet, called Stradanus, Flemish, 1523–1605 Public Domain

Methodology is defined as a system of methods used in a particular area of study or activity. Genealogical methodology goes back thousands of years to those compiled back in ancient times. Confucius' genealogy goes back 83 recorded generations for over 2,500 years. See "Confucius: Ancient family tree." There are other family genealogies that also go back thousands of years. See "8 Oldest Family Trees Ever." There is an old adage that says if it isn't broken don't try to fix it. But how do we know if a methodology is broken and how do we know if it still works?

Let's suppose that I was an actual descendant of someone in one of these old genealogies. What would that really mean? Well, if I were an actual descendant in a royal family such as the Grimaldis in Monaco, I probably wouldn't be living in obscurity in Provo, Utah. Many inexperienced genealogists equate royalty with nobility but there is a real difference. Here is a quote from a website called DifferenceBetween.com article called, "Difference Between Royalty and Nobility."
Royalty refers to the people who are members of the royal family. This includes the king, the queen, the princes, and the princesses. Nobility, on the other hand, is also of high breeding. However, not all nobles are royalty. Nobles can loosely be defined as those who belong to the aristocratic class in the society.
Here is another short description.
Royalty is not something that an individual can achieve. It is an ascribed status. A person has to be born to such a family in order to be royalty. This goes on from one generation to another.  
The short explanation is that if you are royalty you will know it. Nobility is nothing more than aristocracy. But can you be related to royalty? The simple answer is that mathematically speaking, we are all related to royalty because modern genetics trace everyone on earth back to a common ancestor known as the "mitochondrial eve" or the most recent common ancestor of all living humans. See "No, a Mitochondrial “Eve” Is Not the First Female in a Species."

So theoretically, we should be able to identify the relationship of every person on the face of the earth today. Let's assume that we had the DNA of every living person and then calculated the relationship of everyone to everyone else. In essence, we would have a universal family tree of all the humans on earth even if we did not know the identity of every ancestor. Let's further suppose that we took all these DNA relationships and overlaid that information on all the information that is known about historical relationships. The results would arguably establish everyone's relationships and give us a basis for constructing a universal family tree.

If it is theoretically possible to determine how every person alive is related, then it should also be possible to construct that single, universal pedigree of all of the humans on the earth. So far the largest verified family tree is the one created by MyHeritage. See "This May Be the World’s Largest Family Tree Using more than 86 million profiles from Geni.com, researchers created a database that links 13 million people." But FamilySearch.org's Family Tree is specifically intended to be a universal family tree. There may be doubts about its accuracy and the number of duplicates, but there are parts of the FamilySearch Family Tree that are becoming as accurate as possible. If in the future, a way is found to verify the FamilySearch Family Tree with DNA data, it could become the best universal family tree in existence.

New tools require new methodologies. If a universal family tree is theoretically possible, then working with individual pedigrees and genealogies makes no sense. We should all be trying to identify those nodes in the universal family tree that fit our individual family history into the structure otherwise, we are really not doing any "original" family history research, we are just replicating what may already have been done by others. So, unless you start (and continue) your genealogical research with an active investigation of what is already been verified in a universal family tree, you are very likely duplicating the work of others. Once a universal family tree became possible, it became inevitable.

Saturday, March 7, 2020

New Images Feature from FamilySearch opens up unindexed records


The FamilySearch.org website recently added an "Images" search selection to the Search menu (shown above). This new feature is called the Historical Images Tool. The reason for adding this tool is somewhat complicated. Here is an excerpt from the FamilySearch blog post entitled, "Explore Historical Images Tool Unlocks Data in Digital Records," introducing the feature.
Explore Historical Images marks the beginning of a new and different search experience. With this tool, images produced from FamilySearch’s 300+ digital cameras worldwide is made almost instantly available. 
Explore Historical Images helps you navigate to images of historical records that could contain information about your ancestors. Although you aren’t able to search for your ancestor by name directly, you are able to narrow your search by place, date, and other information that was captured when the image was taken.
The reality of the FamilySearch.org website is that, according to the blog post, in 2018, FamilySearch added over 432 million new record images to its online collections. From the January 2020 FamilySearch.org Facts, there are 1.73 billion digital images published only in the FamilySearch.org Catalog. This means that these images are not indexed and cannot be searched in the Historical Record Collections. In fact, there are more images listed only in the Catalog than there are in the Historical Record Collections. Of course, these numbers change constantly but the percentage of records in the catalog will continue to grow at a faster rate than the number in the Historical Record Collections.

Step-by-step instructions about the Historical Images Tool are contained in the above-linked blog post. Notwithstanding the new tool, the bottom line is that the researcher (you) will still have the task of searching through the records one-by-one unless the original records happen to contain some sort of index. The main idea of the new tool is to make users aware of the treasure trove of information that is still locked up in the online images, just one step away from the original paper records or microfilm images.

I still suggest that you may wish to use the main Catalog search that has been available for a long time. You will want to try both for the information you are researching. As an interesting side note, it appears that the probate records my wife and I scanned back in 2018 in the Maryland State Archives are now in the catalog and available for searching.

Friday, March 6, 2020

More than one online family tree? Pros and Cons

Maurice Prendergast, born St. John's, Newfoundland 1858-died New York City 1924
Note: The above image is part of a new online collection of about 2.8 million images put into the public domain on the Smithsonian Open Access website. See https://www.si.edu/openaccess/

One of the tragedies of modern genealogical research is the fragmentation of the huge collections of digital online genealogically valuable resources. There are presently four huge online family tree/database websites and a huge number of smaller more focused collections. But there are also valuable digital resources on websites that do not even mention the words "genealogy" or "family history."

The four large genealogy-focused websites include the following with estimates of the number of images or files or collections or whatever in each. These numbers are calculated differently for each website but will give you an idea of the enormous number of records available in just these four.
  • Ancestry.com -- 32,744 Collections for about 24 billion records
  • FamilySearch.org -- 3.14 billion digital images with 4.93 billion searchable records in 2,724 collections
  • MyHeritage.com -- 11.9 billion historical records in 6,641 collections
  • Findmypast.com -- Over 4 billion searchable records
Why are these numbers a tragedy? Because very few genealogists use all four of these programs when doing their research and either don't know about or ignore the vast number of additional places to look. In addition, even fewer genealogists have a family tree on all four programs.

You may well ask why you would want to have more than one family tree? The first and most common excuse for a second or back-up family tree is fairly common among users of the unified, collaborative FamilySearch.org Family Tree. It is possible that some unresponsible changes to the Family Tree could wipe out sections of your part of the Family Tree and having a back-up of your data and sources makes restoring that information easier. I hear horror stories of irresponsible changes fairly frequently but most of the stories relate to a particular ancestor or at most, a particular line. There are several ways to back up your work but none of them will automatically restore "correct" data to the FamilySearch.org Family Tree. If you have this concern about the FamilySearch.org Family Tree, I suggest you watch the following videos.
Besides back-up issues, there are some other exceptionally good reasons to have a family tree on all four of the major genealogical database programs. The main reason is simple: they all have exceptionally helpful automatic record hints and no, they do not all have the same records. Of course, there are those who claim that these record hints are "frequently" wrong and don't help at all but when I hear that, I often find that the person complaining does not have a family tree on all four programs and does not even try to use the record hints available on the one program they do use. I also find that a significant number of people fail to review the record hints they add to their family tree or in the alternative fail to correct the entries from the information in a validly discovered record. 

I admit that the number of record hints or matches can be overwhelming to some users. For example, I presently have 17,634 record hints waiting to be confirmed and attached for 13,006 records on my Ancestry.com family tree and 2,470 people with 7,737 Record Matches on MyHeritage.com. My own experience is that the accuracy of both is very high. It seems strange to me that a genealogist can be bothered by too much information. It is usually the other way around. 

This past week or so, I helped a patron in the Brigham Young University Family History Library with some Irish research for a "brick wall" ancestor. She said they had been looking for this person for a long time. However, I soon discovered that she was not acquainted with the Findmypast.com website. After making a search for the person she was stuck on, I found his birth record on the parish register where he lived in Ireland on Findmypast.com, Granted this does not always happen, but don't underestimate the huge number of records on just these four websites. 

I think one main reason why more genealogists don't have family trees on all four programs is simple inertia. It takes effort to learn all four programs. It takes more effort to add information to all four trees and harvest the record hints. 

For some genealogists, looking at or participating in a "family tree program" is beneath their dignity. They do research in libraries and archives and do not deign to use such public resources such as an online family tree. Whether through lack of computer skills or for whatever other reason, they can be found teaching an entire class on research in a particular country or writing a book about genealogical research without even mentioning any online sources. Paraphrasing a well-known quote, those who ignore online family tree/database websites will be bound to waste their time looking for sources that are easily found online. 


A New Look for Findmypast


After many years with no significant changes, the Findmypast.com website has had a complete make-over. Here is a summary of the changes from a recent email.
After many years, Findmypast has completely refreshed the site’s design as the brand enters a new decade. While Findmypast’s core features and services remain unchanged, the look and feel of the site has been significantly improved to encourage new users to explore their past while staying true to its roots as the must-have genealogy resource experienced researchers know and love. 
As well as a new color scheme, new record icons and minor navigation changes such as the new ‘Help & more’ button that directs new users to the resources they need, Findmypast is proud to introduce a new brand logo that reflects how family history is unique for each individual.  
Users will see different variations of ‘my’ on the logo throughout their journey with Findmypast. These ‘mys’ have been collected from members of the Findmypast community, so are authentic and personal - just how family history should be. 
Users will see Findmypast’s new tagline - “Where will your past take you?”- used across the site and on social media. This reflects how Findmypast helps researchers across the globe see their bigger picture by learning about their past, present and future. It is a celebration of lives understood backwards but lived forwards. 
Findmypast is also reviewing how they describe their unrivalled collection of records and features by simplifying the language and methods used to speak directly to their members. Findmypast believes family history should be for everyone and seeks to break down the barriers to entry by cultivating a warmer, more encouraging environment for family discoveries. 
These changes are just the latest step in Findmypast’s drive to improve the experiences of all its users. Following on from the release of tree-to-tree hints earlier this year, researchers can expect to see a variety of new features and record collections releases announced throughout 2020.
Fortunately, all of the very useful functions of the website remain with even more extensive genealogically important records. 

Tuesday, March 3, 2020

Is GEDCOM still relevant in today's online world? Part One

Ancient Ruins in the Cañon de Chelle, N.M. (No. 11, Geographical Explorations and Surveys West of the 100th Meridian) Public Domain
Smithsonian American Art Museum and its Renwick Gallery

GEDCOM is an acronym standing for Genealogical Data Communications. "GEDCOM" is copyrighted term and the copyright is owned by FamilySearch. It is an open, de facto specification for exchanging genealogical data between different genealogical software programs. For example, you can download an Ancestry.com family tree in the GEDCOM format. See "Uploading and Downloading Trees."

A GEDCOM file is in plain text. Here is an example of part of one of my downloaded GEDCOM files as it appears in Microsoft Word.

1 DEAT
2 DATE 7 MAY 1699
2 PLAC Aasted, Lindetsborg, Hjorring, Denmark
2 SOUR @S223@
1 BAPL
2 DATE 14 JUN 1910
2 SOUR @S223@
1 ENDL
2 DATE 17 JUN 1910
2 SOUR @S223@
1 FAMS @F247@
1 CHAN
2 DATE 6 DEC 2006
0 @I346@ INDI
1 _UID F2181814B06540AFBCE672FE58D7E56CCC01
1 NAME Jens /Egbertsen/
2 SOUR @S223@
1 SEX M
1 BAPL
2 DATE Submitted
2 SOUR @S223@
1 ENDL
2 DATE Submitted
2 SOUR @S223@
1 FAMS @F249@
1 CHAN
2 DATE 6 DEC 2006
0 @I347@ INDI
1 _UID DEAACB8DD4C745A4B124A8D3C7260731D96C
1 NAME Anne /Pedersen/
2 SOUR @S223@
1 SEX F
1 BAPL
2 DATE Submitted
2 SOUR @S223@

Despite the fact that this appears to be incomprehensible if you have access to the complete GEDCOM standard, in this case, GEDCOM Standard Release 5.5, with some study, you could decode all of this information into what would appear in software programs as entries for individuals in a family tree. The complete description of this particular file is 3,504 pages long in Microsoft Word and has 3,701,054 characters. You can find a copy of the GEDCOM Standard Release 5.5 from 2 January 1996 [Revised 10 January 1996] on RootsWeb.com. See http://homepages.rootsweb.com/~pmcbride/gedcom/55gctoc.htm
Also note the copyright: Copyright © 1987, 1989, 1992, 1993, 1995 by The Church of Jesus Christ of Latter-day Saints. This document may be copied for purposes of review or programming of genealogical software, provided this notice is included. All other rights reserved.

There was a draft revision of the GEDCOM Standard to Version 5.5.1 on 2 October 1999. The very first version of GEDCOM 1.0 was released back in 1984. I am not going to review the politics that propelled the GEDCOM Standard to become a de facto standard in this blog series but I am going to discuss some of the history and the relevance of the Standard today in light of the fact that FamilySearch, after a pause of ten years, released a final version of GEDCOM Version 5.5.1 on 15 November 2019. See The GEDCOM Standard Prepared by the Family History Department of The Church of Jesus Christ of Latter-day Saints and once again, the copyright notice: Copyright © 1987, 1989, 1992, 1993, 1995, 1999, 2019 by The Church of Jesus Christ of Latter-day Saints. This document may be copied for purposes of review or programming of genealogical software, provided this notice is included. All other rights reserved.

It is important to understand that GEDCOM is not a "program" as such. It is a standard for sharing genealogical information between software programs both desktop-oriented and online. Here is a quote from The GEDCOM Standard, page 5 about the content of the Standard.
The GEDCOM Standard is a technical document written for computer programmers, system developers, and technically sophisticated users. It covers the following topics: 
! GEDCOM Data Representation Grammar (see Chapter 1 beginning on page 9)
! Lineage-Linked Grammar (see Chapter 2, beginning on page 19)
! Lineage-Linked GEDCOM Tags (see Appendix A, 83 and Chapter 2, beginning on page 19)
! The Church of Jesus Christ of Latter-day Saints' temple codes (see Appendix B, page 96)
! ANSEL Character Codes (see Chapter 3, beginning on page 77, and Appendix C beginning on
page 97)
It is possible that this release was made as a response to a somewhat earlier release of a GEDCOM 5.5.5 Revision on 2 October 2019 by programmer and genealogist Tamura Jones. See "You waited 20 years, and there's nothing new." I wrote about this release in a post on 30 December 2019. See "New GEDCOM Version 5.5.5 Released" See also the following:

https://www.gedcom.org/

You may also wish to review the BetterGEDCOM wiki at https://archive.fhiso.org/BetterGEDCOM/HOME.html
and the Family History Information Standards Organisation, Inc. or FHISO.

Now, with that introduction, stay tuned for some of my own comments.

Monday, March 2, 2020

MyHeritage Adds Huge Collection of Historical U.S. City Directories

https://blog.myheritage.com/2020/02/myheritage-adds-huge-collection-of-historical-u-s-city-directories/
Quoting from the MyHeritage blog post,
We are pleased to announce the publication of a huge collection of historical U.S. city directories — an effort that has been two years in the making. The collection was produced exclusively by MyHeritage from 25,000 public U.S. city directories published between 1860 and 1960. It comprises 545 million aggregated records that have been consolidated from 1.3 billion records, many of which included similar entries for the same individual. This addition brings the total number of historical records on MyHeritage to 11.9 billion records.
This huge collection of new online records was announced this past week at the annual RootsTech Conference in Salt Lake City, Utah. You need to read the blog post linked above to understand the depth of this mammoth undertaking. Here is a short quote to help you begin to understand this valuable resource.
The city directories in this collection were published by thousands of cities and towns all over the U.S., and each directory is formatted differently. The huge amount of content and its variety made the project more challenging and required the development of special technology to process the city directories. 
We first used Optical Character Recognition (OCR) to convert the scanned images of the directories into text. This process can result in errors in the output, and we created algorithms to detect and correct some of these errors. 
Then, we needed to parse the records to identify the different fields in each record: names, occupations, addresses, and more. The differences in formatting between the books presented an additional challenge. Our team employed methods such as Name Entity Recognition (NER) and Conditional Random Field (CRF) to train an algorithm using a per-book model — meaning that for each of the 25,000 books, we manually labeled a sample of the records and used it to train the algorithm how to parse that directory. Using this model, the algorithm was able to parse the entire book into a structured index of valuable historical information.
One major development implemented in this collection is the consolidation of entries across years of the directories' publications. Here is a further explanation of this process.
After all the information was parsed, we consolidated the records in an unprecedented way. We identified records thought to describe the same individual who lived at one particular address over several years, as published in multiple editions of the city directories. We then consolidated all of those entries into one aggregated record that covers a span of years. This reduced “search engine pollution,” wherein a search for a person would have returned multiple, very similar entries from successive years, obscuring other records. The aggregation makes it easier to spot career changes, approximate marriage dates, re-marriages, and plausible death dates. To our knowledge, the algorithmic deduction of marriage and death events from city directories is unique to MyHeritage.