Some people eat, sleep and chew gum, I do genealogy and write...

Friday, June 13, 2014

What ever happened to the genealogical data standards issue?

I could answer this question in one word. APIs. The emphasis in transferring information between genealogy programs generally has been all but superseded by the large online genealogy database programs efforts to integrate their data with FamilySearch. However, the issue of exchanging information between different individual genealogical database programs remains unaddressed and unsolved. There are a few individuals who keep addressing the issue, such as Tony Proctor with his Parallax View blog, but the rest of the online genealogical community is strangely silent on the subject.

The main issue, adequately transferring complete genealogical data from one program to another still remains. But it appears that other than the ongoing GEDCOM X effort, little or no progress has been made since meeting held in 2013 at RootsTech. Meanwhile, almost all of the currently available desktop oriented genealogy database programs are still dependent on the old GEDCOM standard.

The effect of a lack of a standard way to transfer data is that as the genealogy programs become more complex and the demands for recognizing more data fields continues to grow, there is more and more lost information as data is transferred from program to program. This is especially true as more programs ignore the old GEDCOM standards or face the old standard's limitations.

Last year, 2013, there was a detectible movement in the Family History Information Standards Organisation (FHISO), but since 30 July 2013, almost one year now, there has not been anything posted to their website. Likewise, its predecessor organization, BetterGEDCOM, has shown no activity.

Now there are many ways to approach this issue:

  1. Drop the issue and forget about standards.
  2. Write sarcastic blog posts.
  3. Start pushing for some kind of data transfer standards
  4. Wake up and notice what is actually going on

Well, it turns out the fourth option is the one to consider. What is happening is that the four very large genealogy database companies (VLGDCs) are actively working on the problem of sharing data using APIs (Application Programming Interfaces). FamilySearch has data sharing arrangements with each of the three other VLGDCs. I do not pretend to have the slightest idea how they are sharing data but databases (collections) on FamilySearch.org are showing up on the three other companies; Ancestry.com, MyHeritage.com and possibly, findmypast.com. In addition, at least with Ancestry.com presently, there is a limited amount of data sharing between family trees on the two programs. It seems inevitable that this sharing arrangement will sometime include images and other media files. Three of these large companies, Ancestry.com, FamilySearch.org and MyHeritage.com, have desktop programs that can exchange data (i.e. synchronize) files to a greater or lesser extent with a desktop genealogy database program. In the case of Ancestry.com and MyHeritage.com, it is their own proprietary programs. In the case of FamilySearch.org's Family Tree, there are several third-party programs that can move data (share data) back and forth with the online Family Tree.

There is presently no way to test how effective it would be to try to move a complete online tree from Ancestry.com to Family Tree and then to a desktop program. Likely this would be accomplished by moving separate data fields either one-by-one or in some kind of batch process.

The danger here is that FamilySearch.org's Family Tree may become over-saturated with the messy files on the other programs.

So while the genealogical community essentially ignores the issue the VLGDCs are in the process of programming what will likely become a de facto standard. This may be accomplished by driving any unconnected software company out of business.

I am not asking anyone to do anything. I am just commenting on the evidence.


Thursday, June 12, 2014

Comments on old news -- Ancestry.com Focuses on Core Offerings

One side effect of being "offline" for essentially three weeks is the fact that I find out about stuff much after the initial announcement. This applies to a blog post by Ancestry.com entitled, "Ancestry.com Focuses on Core Offering." This is not much of surprise since I have been writing about the demise of one of the programs for quite a while. The five programs being discontinued on September 5, 2014 include the following;
I have been writing about Mundia.com for a long time. See my post of February 18, 2014 entitled, "Is Mundia.com Dysfunctional?" for an example.

Ancestry.com does make an exception for Genealogy.com and explains (or not):
Genealogy.com is the exception to the rule, and will continue in a slightly different form. If you are an active member or subscriber to one of these services, you will be contacted directly with details of how to transition the information you’ve created using these services.
More information about the changes is on the Genealogy.com FAQs page.  There the explanation is a little more detailed:
Over the years we have built up a variety of products that enable our users to discover, preserve and share their family history. We recognize that there are a lot of ways that we, as a company, can make family history easier, more accessible and more fun for people all over the world.

In order to do this, we need to focus on our core offerings to ensure we're delivering the best service and best product experience to our customers. To that end, we've decided to invest more aggressively in our core Ancestry.com business and plan to retire all member activity on Genealogy.com as well as portions of the Genealogy.com service.
Current users of Genealogy.com have been given the following options for preserving their data on the website:
You will have the ability to print or export My Trees and manually print or save additional content you've uploaded until September 5, 2014. To preserve the information you've added, please log in to your account to export, print, or save your information.
This scenario is a wake up call to all of the genealogical community. Do not ever rely on one method, even an online method, to backup your genealogical data. What is happening there is reality. Since I came through the 1960s, the mantra of my generation still holds true in most cases: Trust No One Over 30. I can say that with certainty to any large commercial enterprise online or otherwise. By the way, almost all of the online genealogy companies, with a few notable exceptions, are less than 30 years old so I would have to modify the original call and shorten it to "Trust No Online Data Storage" and always rely on more than one storage media.

Growing your genealogical tree

Alaskan Tundra in Danali National Park
Whatever the motivation, all of us who are pursuing our family history or genealogy, begin with gathering information about our closest relatives. I find that this interest usually extends to parents and grandparents. In some instances, the motivation to keep gathering information comes from the desire to connect with more remote relatives, perhaps those who emigrated from another country, fought in a war or became famous. Some of the reasons for this interest might include stories passed down to us through family traditions, surname books, heirlooms and other artifacts inherited from our family. It is inevitable that as we pursue this initial interest, we will reach the point where obtaining more information becomes increasingly difficult. What do we need to do to move forward and obtain even more information about our ancestors?

This is one of the key obstacles faced by every genealogist at every level of expertise. The answer seems simple but, in fact, involves some very complicated processes that must be learned. Some genealogists face this obstacle by trial and error as I did in my early learning process. But in our network-connected world, many more turn to the simple solution of copying the work of others either from a book or an online family tree. 

I have been traveling through the Alaskan forests recently and there are areas where the trees grow so closely together that they become impenetrable. They call these closed forests. These thickets of trees become such a barrier that even the bears and moose cannot force a way through. Imagine the difficulty of trying to identify or work with one tree in the midst of that tangle. I fear that this is where many of the world’s potential family historians find themselves stopped. I further fear that many of them do not even realize they are entangled in such a barrier.

Those working on the Alaskan highways have a rather simple solution to the problem of the dense forest along the roadways; they cut down all the trees back from the roadway to help prevent forest fires and minimize the danger from animals walking onto the highways. In short, they start over.

In this analogy, I see a way for us to proceed with growing our own family trees. Sometimes we need to start over. I don’t mean that we need to be like the Alaskan Highway Department and throw it all out (although that may be a good option for some), what I do mean is that we need to begin again with gathering information about our nearest relatives and systematically do the same thing, working our way through our family tree. This is most appropriate for those who “solved the problem” by copying information from others rather than verifying the information through their own research. If the copies records contain source citations, perhaps the task will be made somewhat easier, but in many cases, the copied records are bereft of citations and the effort must include documentation.

It is my observation that only a very small percentage of those seeking information about their families will take this step and really begin to become genealogists and not just name collectors. You might want to examine your own family tree, whether it is online or on paper in your files. Do you have source citations for each of your ancestors and their family members? If you do not have source citations, how do you know if the information you have recorded is correct? Where would you go to verify the information previously collected? 

I saw another thing in Alaska that reminded me of this same issue with genealogical research. Many of the trees growing the in the vast Alaskan forests were stunted and deformed. I learned that these trees were growing on permafrost. There was only a very shallow layer of ground that thawed out each year where the trees could grow and so the trees had shallow roots and were liable to be blown over in a storm. Some of the trees were hundreds of years old but were only a few feet tall and had only grown to a few inches in diameter. Our lack of sources is like that permafrost layer under the trees in Alaska. We can't get our family tree to grow because we have created a barrier where further research has stopped for a lack of source citations that support further growth and like the Alaskan trees, our research is likely to be blown over when questioned because there are no deep supporting roots. 

OK, enough of analogies. The simple question is this: how much documentation have you collected and recorded about yourself and your immediate family? That is a good place to start. 

Wednesday, June 11, 2014

How will pushing data affect genealogy and genealogists?

There is a rapidly exploding trend in online genealogy to “push” data to genealogist by means of online database and family tree programs. Personally, I am a active proponent of such “pushed data,” essentially for the reason that it automates and accelerates “routine” genealogical research. But after nearly 40 years of trial experience as an attorney, I am well practiced in examining all sides of any issue, particularly if that issue tends to be controversial. For this reason alone, I have analyzed some of the negative consequences of “pushed data.” 

One side note. Controversy in genealogy is about as mild and inoffensive as anything that might be called controversy. Sometimes I hit on an issue that raises the average hits on my blog somewhat, but I use the word "controversy" advisably. The reason I use the word here in reference to pushing data is that there is substantial confusion among the genealogists I talk to concerning both the need for such a system and the effects that the proffered "sources" have on researchers.

Before I go much further, I need to answer the question of what is pushed data? 

Many genealogists use online database programs to research collections of digitized source documents and indexes of source documents. There are likely thousands of such websites scattered around the world, from small collections of a few specialized documents to web based mega-data providers that have millions and even billions of searchable records. These websites are both fee based and free depending on the provider's motivation. As long as those records remained passively supplied, researching the records was an almost exact analog of researching traditional paper-based or microfilmed records only much more convenient. However, at some point the purveyors of these online collections began proactively pushing their records to the users. This was first done in the realm of connecting different nodes on online family trees. For example, if two users shared the same remote ancestor, the programs began telling the two users of the potential connection.

Suggested family tree connections rapidly evolved into the databases suggesting potential data sources. I have noted in other blog posts that the technology and programming behind these systems of automated source suggestion are extremely sophisticated and complex. It may seem trivial to match your great-grandfather or mother to a corresponding record in a database, particularly if that database has been transcribed and indexed, but in fact, this is an almost monumental task. The reason for this difficulty should be clear to any genealogist who has spent a significant amount of time researching original records. It is one thing to find the record, it is quite another thing to read and interpret the record accurately. Original records are sometimes vague, indistinct, inaccurate and illegible. Writing computer programs that can compensate for these limitations and perform consistently with positive results, is a monumental challenge.

Notwithstanding this difficulty, many of the online database companies have begun the process of matching user’s ancestors to source records, either through automated programs or with user intervention. Some of the online programs do a phenomenally good job of finding sources. This source-finding ability appears to be a natural outgrowth of the user oriented search engine technology used by the websites to help users find indexed records. Actually, it is quite a different technology. It is not just a better search engine, it is revolutionary way of looking at genealogical records and matching those records to the right person through a set of in depth algorithms that consider entire pedigree segments rather than the name, date and place of some individual. Some of the previously available "Advanced Search" capabilities previewed this type of technology, where the user could add the names of some selected relatives to assist in identifying the target ancestor.

In the competitive online world of commercial genealogy programs, why not develop technology that that will not only differentiate similar products but also become a value added attractor to new users? One challenge for the companies developing this new technology is the difficulty in communicating the benefits of the matching technology to naive users. This difficulty is dramatically illustrated by how few of those who put their family trees online are further motivated to take advantage of the automatic or semi-automatic source and tree matching functions. As I stated, I believe there is a substantial benefit to the genealogical community, but those benefits do not seem to be obvious to most of the people who have family trees online.

So what are the drawbacks of pushed sources? I would suggest that the first and most serious drawback is that the typical user has no idea what to do with the suggested source or why such a source might even be useful or necessary. So far, none of the online programs suggesting sources have provided either motivation or support for the process. Now, what I mean by this lack of support is not that the process of adding sources is not adequately explained, but the rationale for adding sources is all but ignored. It is not the mechanics of adding the sources that is the challenge. The real challenge is convincing those who post their family trees online that sources supporting the information in the family tree are necessary. To a seasoned genealogical researcher, the idea that any fact or event recorded in a family tree should be supported by a source citation is elementary. But in the real world of online family trees, this is far from reality. I am purposely avoiding using any particular online program as an example, but it is clear from an examination of even a few user submitted online family trees that sources citations are not a high priority.

Even if a family tree submitter is convinced, for whatever reason, that sources are important, the ease of obtaining those sources is a trap. If I were to upload my family to one of these websites that push sources and immediately got a suggested U.S. Census record, how would I know what to do with the record. Does the average user of a family tree program even know what a census records is? Why would I think that even more sources might be helpful or necessary? In other words, these programs are providing the end product of researching a record without having a researcher to evaluate the information contained in the record and integrate that information into a coherent ancestral record. So I have a source, so what? That is the question that needs to be answered. If the users of these programs do not see the need for sources in the first instance, why would they spend the time to look at sources and examine them critically? How would they know that that they needed to look elsewhere for additional information? 

My concern is that by offering what are essentially "fast food" sources provided with little or no effort on the part of the user, the programs are becoming a disincentive to further useful research. If you got dessert all day, why would you want meat and potatoes?

From another standpoint there is also an issue with the sources themselves. Even granting the online websites great technical skill and accuracy in connecting the source to the right person, the issue becomes the source. What if the information in the source is inaccurate or incomplete? How is the naive user supposed to know what to do with this misleading source? 

I think that adding automatic source matching to the online programs makes them immensely more useful to seasoned, well-founded, genealogical researchers. But there is one last issue. The experienced genealogists are also unimpressed with the matching technology because they don't know how to use it. They are overwhelmed with the number of sources offered and resent the programs for detracting them from their research goals. 

It looks like to me that the matching programs have a long way to go before they are generally beneficial to the genealogical community. As the number of online sources aggregated to these programs increases, they become that much more valuable to the prepared researchers, but at the same time, they become a stumbling block to those who are not well founded in research principles. 

By the way, from my contacts with the large genealogy companies, I am reasonably aware that those developing these technologies are aware of the problems and challenges and are working to overcome those same issues. I am very positive about the future and see this technology becoming a tremendously time-saving tool. 





Dealing with a multitude of source suggestions

There are currently a number of programs that make source suggestions and connections to other family trees to their subscribers within the context of hosted family trees. They include, at least, the following websites:
There are probably others and I understand that FamilySearch.org Family Tree may also start providing such automatic suggested links to sources. Right now, FamilySearch.org's Family Tree will search for sources but the user has to instigate the search. However, in the descendancy mode, Family Tree will suggest sources for selected individuals. Because of the nature of FamilySearch Family Tree, there is no need to suggest connections to other family trees because the program is a unified family tree. Although Mocavo.com hosts family trees, it suggests sources only, so far. But it does suggest that you join a "surname" group for individual connections. It should also be noted that Geni.com is owned by MyHeritage.com and uses some of the same technology.

The programming for both processes, connecting family trees and finding sources for individual ancestors, is extremely complicated. Variations in the way individuals are recorded make determining connections difficult. This process can be a challenge even for an individual human researcher.

As far as I am able to determine, Ancestry.com was the first to automate the process. A significant advance in accuracy was released by MyHeritage.com in 2013. Because of the complexity of this type of matching, many of the programs are surfeited with "false positives." This means that a suggested connection or source is attributed to the "wrong" person. For example, your ancestor lived in New York and the sources or connections are to people living in other states or even other countries.

Because of these false positives, there are recurrent questions about this practice that fall into two general categories. First, the questions involve the accuracy of the suggested sources and second, they involve the number of suggestions.

If you are not familiar with this particular type of feature of any of the above programs, you should be. By and large, these types of automated source and family tree connections are extremely valuable aids to research. Each of the programs listed above (with the exception of FamilySearch Family Tree) host multiple (millions of) individually submitted family trees.

If you have a family tree on any one of the above programs, you are likely receiving notifications of suggested "matches" to sources and to other family trees. In some cases, the number of these notifications can be overwhelming. For example, right now on my Ancestry.com family tree, I have only 326 suggested sources these include connections to other user submitted family trees. However, this number has been over 1000. On my family tree on MyHeritage.com, I have 501 Smart Matches to other family trees pending and 7096 pending Record Matches to sources. In addition, if I click on any of the Record Matches, I will have many more suggested sources.

The accuracy of these particular programs, Ancestry.com and MyHeritage.com are very high. It is very difficult to compare the two programs because they have different sets of source documents and the suggestions do not overlap except when they both have the same source.

Depending on the individuals in your own family tree and whether or not the source documents in any of these programs apply to your particular ancestry, you may or may not have this type of response to the automated search and tree matching functions. But in many cases, the response to your family tree is overwhelming. How do you deal with these huge numbers of sources and connections?

My answer is simple, ignore them. There is no reason to feel anxiety over the fact that there are so many sources and connections available. Use the programs as a tool for your own research objectives. If you have the objective to add sources to all the individuals in your family tree file, then do so. Work through the suggestions systematically but realize that the programs will always suggest more than you can handle. On the other hand, if you wish, you can focus on individuals in your family tree and if there are suggestions they might be helpful, but if not, then you are in a waiting mode. Remember that these programs all add new collections and sources on a very regular basis. You may not have a suggested source for an individual today, but that could change tomorrow.

What if you want to use these programs to do your own research and not use their automated search and matching? You can, of course, do this but then you are basically wasting your money. These programs are designed to do the work for you. If you do your own searches, you are not using the full potential of the programs. You must have uploaded a family tree and you need to allow the programs to make full use of their potential.

Monday, June 9, 2014

Towards Defining Ownership in Genealogy

Ron Tanner of FamilySearch often speaks about “MyTreeistis” or claiming ownership of an online family tree. I commonly remind those whom I teach that we do not own our ancestors and that slavery is not longer legal in the United States and most other countries. But seriously, the question for genealogists is what do we “own” and what is not owned?

Ownership implies some sort of control by the individual “owner,” sometimes, exclusive.  But ownership is a both a legal and a cultural concept that varies considerably from culture to culture and country to country around the world.

Let me start this post by asking this question: When the Mayflower passengers landed in Massachusetts Bay in 1620 who owned America and particularly, who owned the land where they set up their first colony? It is interesting that these “colonists” were given ownership of land in amounts depending, in part, on when they first arrived in America. If you know any history, England claimed ownership of the American land, more particularly the King of England. But it is also equally well known that other European countries also claimed ownership to some of the same parts of America. Those countries included Holland, France and Spain. Wasn’t this claim of ownership entirely based on a cultural and a legal disregard for the claims of the people who already “owned” the land? So, essentially, occupation of the land by an Englishman created an “ownership” interest while prior occupation by the Indians (Native Americans if you like) conferred no such interest? It is true that in some instances the encroaching Europeans made token payment to the inhabitants for ownership of the land. But it is equally true that the Amerinds did not understand the concept of ownership by a foreign king they had never seen nor heard. In addition, it was the Europeans who set the price.

Now what does this have to do with genealogy? Quite a bit actually. As we do research about our ancestors, we are like the Europeans landing in America. What we research and what we claim as our own, is already owned by the people who are there before we do our research. In other words, we are exactly like the European settlers. We claim ownership based on mere possession and totally disregard the ownership rights (if any) of all of the other members of our family who are equally related to those same ancestors. In doing this, we are really squatters not owners.

Now, you say, but what about work product, copyright and attribution? Aren’t we the actual owners of our own “genealogy?” Apparently, this belief comes from a quasi-real concept of genealogical homesteading. If we find the information and work with it for a set number of years, we can assume that we own it. Isn’t that correct? Actually this is not correct. There is no such process for acquiring ownership, either cultural or legal. So, is there any part of our genealogical research to which we can claim ownership?

Before answering that particular question, I need to ask why the question arises in the first place? As genealogists, why do we think we need to have ownership of our research? Our Western European society, heavily influenced by English Common Law, provides for an individual’s ownership of both real and personal property subject to some important restrictions. For example, if you own your own home, is your ownership absolute? What would happen to your ownership if you failed to pay your taxes? If you “own your own home” then you must realize that if you fail to pay your taxes, then the city/county/state etc. can impose a “lien” on your property and after continued failure to pay, take possession of and sell your property to satisfy that lien. Aren’t you in effect “renting” your property, through the payment of taxes? Do you really own anything? In a strictly legal sense, we do not actually own anything. Everything we own is claimed by some government or another and could be taken from up by any claim of necessity or right by those governments.

I am not ignoring the greater religious or philosophical questions of “ownership” and how we all die and take nothing but our experiences with us and this may also apply to our genealogical research.

Without getting into a discussion of different religious, social and cultural concepts of ownership, suffice it to say that what we believe we own is based on our beliefs and our culture. For example, going back to the native population of America at the time of the first European colonization efforts, the Indians did not have the same concept of ownership as the Europeans (obviously). The ideas of individual land ownership were largely absent. Ideas about personal property ownership varied greatly from place to place. But all these belief systems of the Indians gave way to the imposition of European law and society.

If we think about this idea of ownership for a while, as I mentioned previously, we will begin to realize that ownership is entirely a culturally and legally based concept. We do not individually determine what we own and what we do not own. These decisions are already made for us by our society at large.  Any claim to private ownership is merely an illusion. Likewise, claims to ownership of our genealogy is also an illusion.

In our society in the United States, we recognize different levels of ownership. For example, land can be owned outright (fee simple) or rented. Unfortunately, when we start talking about research or writing instead of real property, the terminology and concepts get a little bit slippery. Original work may or may not be covered by the U.S. Copyright laws. In addition, we have strong moral imperatives such as work product and attribution. We can’t ignore any of these cultural artifacts that impose rights and limitations on our property, but once again, do any these issue apply to genealogical research? To answer this, I need to resort to a series of hypothetical situations.

Let’s suppose that I go onto one of the many online genealogical resource programs and compile a pedigree chart with supporting documentation. At this point, the information I have compiled consists entirely of names, dates and places with supporting documents. Assuming for this purpose that there are no copyright claims to any of the documents, a researcher’s compilation adds nothing that makes the work subject to copyright law. But what about attribution and work product? Clearly, the compilation becomes a personal work product. If someone copies the work, they should give the researcher credit for his or her efforts. However, neither the claim of work product nor the moral obligation to provide attribution create an ownership interest in the corpus of compiled genealogy. A researcher may be very upset when someone “copies” their research, but that feeling arises because of the researcher’s acculturation and not as a result of any right of absolute ownership.

Let’s further suppose that the researcher adds original comments, insights, etc. to the research. These additional original works may or may not be subject to copyright. There is no way to automatically determine at what point a fact becomes an original work. This particular function is handled on a case-by-case basis by our legal structure pertaining to copyright. Separating the “original” portions of the work from those parts that are not subject to copyright claims is extremely difficult. These issues become almost insurmountably difficult when you consider the fact that merely posting the information online on some of the various family tree programs available may seriously affect your rights to claim copyright protection. In many cases, posting the information online involves giving a license to the hosting website to the content. In other words, the researcher then transfers part of his or her ownership to the hosting website.

But here is the reality of the situation. If you publish, copy and provide, post online or do anything with your research, anyone in your family has equal claim to the resultant facts. Even if they copy those portions of “your research” that you claim to be original, the analogy that has been used in the past is like casting a bag of feathers to the wind. How are you going to enforce a copyright claim? Likewise, becoming obsessed with the issues of work product and attribution are meaningless in the context of ownership. Do you believe that someone will subsequently pay your for your work? If so, make sure you put the work into a format that can be sold as product. In other words, formally publish your work in a surname book (either ebook or on paper) if you want to “sell” the work. Very, very few of this type of publication even garners enough sales to pay for the cost of publication.

So, when you start to have obsessive thoughts about your genealogical research and begin to become defensive about others in your greater human family using or copying the information, chill out. Take a moment to reflect on the realities of ownership. I am aware of some people that are so obsessive about their genealogical research that they won’t even allow anyone to view it because they think it will be copied. This attitude stops working the moment that person dies and his or her heirs chuck the whole pile of paper in the nearest dumpster.

I fully realize that there is nothing much I can say or do to change the attitude of genealogists towards their claims of ownership to their genealogy. But you can’t fault me for trying. I am always thankful for comments, both positive and negative. But before you tell me how you are certainly entitled to your copyright claims, think that through. What are you really saying? Do you really intend to spend thousands or tens of thousands of dollars defending your work from copyists?

Sunday, June 8, 2014

Online Digital Newspaper Collections by State -- The Lists Introduced

There are two preliminary parts to this blog post which include an introduction and a review of the applicable copyright law. Here are the links should you care to review the background and issues of this very interesting topic.
There are quite a few collections of newspapers that cover extensive blocks of time and geography, i.e. they include more than one state's newspapers. It is important to understand that there is some overlap between these huge online collections, but any thorough genealogical search would necessarily require searching every single collection. Of course, that could become a problem since most of these collections are subscription based and not only does the genealogist have to find all of the collections, they also have to figure out how to access them and possibly pay for the content. I say this so that the potential researcher does not feel comfortable ignoring the subscription based sites and only researching the free online content. 

You might also recognize that the effectiveness of the various search engines and the degree to which the optical character recognition programs work affects the ability of a researcher to find specific content using a search. There is really no way that a careful researcher can be assured that there are not important facts about any given ancestor other than to do a page-by-page search, assuming that the online project provides access to multiple pages of the same search. Researchers should also recognize the fact that helpful information may be contained in paid advertising and display advertising in the newspaper digitization project may not have been included, especially if the advertising consisted of images rather than text. But if you come from an old genealogical tradition, you are used to searching microfilm page and page and this is no different.

As it turns out, unlike digital maps websites, there are exhaustive online references to newspaper collections listing each state of the United States in detail. It also turns out that there are a huge number of websites, far more than you could imagine. There are hundreds of websites. Just think what a great opportunity this is. You will never run out of research opportunities.

There are a substantial online lists of available digitized online newspaper collections. See the following websites for lists:
Here is a list of the multi-state online digital newspaper projects that I have found. I do not pretend that this is an exhaustive list, because these collections are sometimes hard to find online and also because new projects pop up frequently. Just because I was unable to find a specific newspaper project for any of the states or territories does not mean that there are no online digitized newspapers from that jurisdiction, any such content may be included in one or more of the large collections. 

Now on to the state-by-state list. I am listing the states but still have very limited access to the Web here in Alaska and will republish this shortly with the data. Thanks for your patience.