My notes are going out raw, so excuse any issues like making sense or not.
Jay Verkler,
Elements of a Community Record
Digitally born records, those that were created on a computer directly. Identifying family relationships in various records across different databases.
This is a more expansive view of the types of information being stored online about an individual or family.
The final element of this global way of looking at records. Creating a source authority with links directly to the archive.
You can likely review this entire presentation on a shared community of records. Jay puts in a plug for having all of this in an open source system and not a proprietary system. FamilySearch has a number of very large and influential partners in this new open system. brightsolid, Archives.com, Geni, Ancestry.com, FGS, to state few.
I strongly suggest that you review this presentation if at all possible. The presentation closes with a very entertaining sketch showing how records and research can be collaborative across the Internet in a variety of applications across the world.
Only the begging, more to come
Thursday, February 2, 2012
RootsTech First Keynote Address - Jay Verkler
Jay Verkler is the outgoing CEO of FamilySearch and is being replaced by Dennis Brimhall. All of the FamilySearch employees and outstanding credentials. Jay is a charismatic presenter and is extremely knowledgeable about the technology which is the theme of this Conference.
As you can probably guess, I can't listen and type at the same time, but I will try to give you all a flavor of the presentation. By the way, this presentation is being carried live on the RootsTech website. They provide us Bloggers with a front row seat for the keynotes, so we can see what kind of makeup they use (just kidding). Even though I am deaf, I probably will have no problem hearing. It is quite loud.
The first speaker is David Bruburgger, Senior VP of Engineering who introduces Dennis Brimhall, the new CEO. He is new to genealogy but not new to leading large non-profits. Dennis Brimhall in turn introduces Jay Verkler. Jay Verkler, who has been CEO for the last ten years, was the one who moved FamilySearch into the technology world.
Jay reviews the experience of the Conference in 2011 when they had over 3000 in attendance. This year there are over 4300 registrations. He then takes a leap to talk about the future in 2060. Points from the talk including some of the new "buzz" words used:
A Genealogical Conclusion: capturing who the person really is with all of the existing relationships.
One example is the Facebook Timeline view. (Facebook shows every post you have ever made chronologically).
Demonstration of a person page that is interactive and shows the details about an individual, packaging the important parts of a person's life.
What does it take to do this?
Community Framework for preserving and interchanging information
FamilySearch working on new GEDCOM, same standard, multiple uses
Jay introduces Robert Gardner and Dave Barney from Google. Yes, Google is here also. I notice they are using a MacBook Pro. Just a comment on technology. Dave talks about Genealogy Microdata Tags.
Links will disappear. What changes need to be made? Permanent links? A solvable problem.
Need for "Authorities" which provide the rules for search engines to recognize records. Getting companies to use the same set of Authorities. More powerful search engines.
As you can probably guess, I can't listen and type at the same time, but I will try to give you all a flavor of the presentation. By the way, this presentation is being carried live on the RootsTech website. They provide us Bloggers with a front row seat for the keynotes, so we can see what kind of makeup they use (just kidding). Even though I am deaf, I probably will have no problem hearing. It is quite loud.
The first speaker is David Bruburgger, Senior VP of Engineering who introduces Dennis Brimhall, the new CEO. He is new to genealogy but not new to leading large non-profits. Dennis Brimhall in turn introduces Jay Verkler. Jay Verkler, who has been CEO for the last ten years, was the one who moved FamilySearch into the technology world.
Jay reviews the experience of the Conference in 2011 when they had over 3000 in attendance. This year there are over 4300 registrations. He then takes a leap to talk about the future in 2060. Points from the talk including some of the new "buzz" words used:
- 7 billion people online in 2060
- Proportionally more people will be interested in genealogy
A Genealogical Conclusion: capturing who the person really is with all of the existing relationships.
One example is the Facebook Timeline view. (Facebook shows every post you have ever made chronologically).
Demonstration of a person page that is interactive and shows the details about an individual, packaging the important parts of a person's life.
What does it take to do this?
Community Framework for preserving and interchanging information
FamilySearch working on new GEDCOM, same standard, multiple uses
- Interchanging between applications and between those records online API standard
- Links and linked information
- Embedded media
- Clear model
- RESTful interfaces
- Extensible
- Common types and sturctures
Jay introduces Robert Gardner and Dave Barney from Google. Yes, Google is here also. I notice they are using a MacBook Pro. Just a comment on technology. Dave talks about Genealogy Microdata Tags.
- Microdata
- Schema.org
- Historical-data.org
Links will disappear. What changes need to be made? Permanent links? A solvable problem.
Need for "Authorities" which provide the rules for search engines to recognize records. Getting companies to use the same set of Authorities. More powerful search engines.
Onsite at RootsTech 2012
You can expect quite a few posts next two days. If you want to keep
current with RootsTech, just check-in occasionally and also search the
other blogs. Besides the official RootsTech Bloggers there are dozens of
others, so there will be plenty of coverage.
Our first keynote speaker is Jay Verkler, most recently CEO of FamilySearch. One thing I have noticed that is different from last year is that Brigham Young University has a lower profile. Last year they were much more visible. One very enjoyable thing about this conference is meeting and greeting bloggers from all over the world and also putting faces with those who have been sending me emails.
It is a rainy and cold day in Salt Lake, but fortunately I did not have to drive this morning and the entrance to the hotel is only about 50 yards or so from the north door of the Salt Palace where the Conference is being held. This is all so overwhelming to a small town boy! Wait, I'm from Phoenix, the sixth largest city in the U.S. and a 37 year veteran trial attorney, nothing, but nothing is overwhelming. But it is very nice and pretty noisy at the moment.
Well, I see Jay Verkler going towards the stage. Time to post this and start another.
Our first keynote speaker is Jay Verkler, most recently CEO of FamilySearch. One thing I have noticed that is different from last year is that Brigham Young University has a lower profile. Last year they were much more visible. One very enjoyable thing about this conference is meeting and greeting bloggers from all over the world and also putting faces with those who have been sending me emails.
It is a rainy and cold day in Salt Lake, but fortunately I did not have to drive this morning and the entrance to the hotel is only about 50 yards or so from the north door of the Salt Palace where the Conference is being held. This is all so overwhelming to a small town boy! Wait, I'm from Phoenix, the sixth largest city in the U.S. and a 37 year veteran trial attorney, nothing, but nothing is overwhelming. But it is very nice and pretty noisy at the moment.
Well, I see Jay Verkler going towards the stage. Time to post this and start another.
What is the SSDI? And why do I care?
One of the first and most valuable resources used by beginning
genealogists in their research in the United States is often the Social
Security Death Index or SSDI. On the old FamilySearch.org
website, the
SSDI was one of the very few sources that could be considered original.
Nearly all of the other records on this valuable site were user
contributed family trees or extracted records. Many researchers got
their first real lead on a family member by locating his or her death
date and location on the SSDI. The SSDI also listed the dead person's
Social Security Number allowing the research to order, for a fee, the
original Social Security Application Form.
Even very experienced researchers found the SSDI a quick way to locate people who lived too recently to show up in any other records due to modern privacy and buerocratic limitations on records.
The value of this extensive listing extends way beyond the realm of genealogy however. For example, the SSDI is used by insurance companies to check whether or not a death claim is valid. If an insurance beneficiary's claim is not substantiated by a corresponding record in the SSDI, then there is a reason to investigate further. Attorney's use the SSDI to find out if someone they are searching for is deceased.
Presently, there is a bill pending in the U.S. Legislature that would seriously limit the use and value of the SSDI. The proponents of the bill are reacting, I believe inappropriately, to a situation where dishonest taxpayers are falsely claiming dependents by using an unrelated recently deceased child's Social Security Number. This problem is not an issue with the SSDI, it is a problem with the way the IRS handles tax returns. In other words, this is not an identity theft issue, it is a tax issue. A deceased child's Social Security Number is associated with the numbers of his or her parents. The IRS could simply verify that a child's Social Security Number matched the parents' number and the problem would be solved. Rather than requiring this simple step, politicians want to use the emotionalism of the loss of a child and the use by another person of the dead child's Social Security Number as a springboard to make a valuable genealogical record even more difficult to use and may destroy the use of the SSDI altogether.
Stay tuned for more specific information about this serious issue and what you can do to help. Go to the website of the House Ways and Means Committee for information about hearings that are being held right now about this bill.
Even very experienced researchers found the SSDI a quick way to locate people who lived too recently to show up in any other records due to modern privacy and buerocratic limitations on records.
The value of this extensive listing extends way beyond the realm of genealogy however. For example, the SSDI is used by insurance companies to check whether or not a death claim is valid. If an insurance beneficiary's claim is not substantiated by a corresponding record in the SSDI, then there is a reason to investigate further. Attorney's use the SSDI to find out if someone they are searching for is deceased.
Presently, there is a bill pending in the U.S. Legislature that would seriously limit the use and value of the SSDI. The proponents of the bill are reacting, I believe inappropriately, to a situation where dishonest taxpayers are falsely claiming dependents by using an unrelated recently deceased child's Social Security Number. This problem is not an issue with the SSDI, it is a problem with the way the IRS handles tax returns. In other words, this is not an identity theft issue, it is a tax issue. A deceased child's Social Security Number is associated with the numbers of his or her parents. The IRS could simply verify that a child's Social Security Number matched the parents' number and the problem would be solved. Rather than requiring this simple step, politicians want to use the emotionalism of the loss of a child and the use by another person of the dead child's Social Security Number as a springboard to make a valuable genealogical record even more difficult to use and may destroy the use of the SSDI altogether.
Stay tuned for more specific information about this serious issue and what you can do to help. Go to the website of the House Ways and Means Committee for information about hearings that are being held right now about this bill.
Wednesday, February 1, 2012
Call to Action
Last year the Bloggers' dinner hosted by FamilySearch was a quiet,
dare I say, contemplative affair with lengthy conversations about the
world of genealogy blogging. This year was a total contrast. The dinner
and conversation were secondary to the important issues of the day. The
first, and very important issue is the introduction of the 1940 U.S.
Census on April 2, 2012. Digitized copies of the Census, not just scans
of an old microfilm record, will be made available by the National Archives. Immediately, the 1940 U.S. Census Community Project, a joint initiative between Archives.com, FamilySearch,
findmypast.com, and other leading
genealogy organizations, will
coordinate efforts to provide quick access to these digital images and
start indexing these records to make them searchable online
with free and open access. This is a call to action. The indexing of the
1940 U.S. Census will be handled by the FamilySearch Indexing
website and program. The more people that volunteer, the faster the
Census will be indexed. You will hear a lot more about this shortly.
Now, on to the other important issues. The U.S. Congress is right now contemplating the passage of two pieces of legislation, one concerning the Social Security Death Index (SSDI) and the other the Model Vital Statistics Act. Passage of either or both of these bills could seriously and very detrimentally affect the entire genealogical community. I will have a lot more about these bills in the immediate future, but right now, information about the bills can be obtained from the House Ways and Means Committee website.
The changes to the SSDI could totally eliminate this valuable resource. The Model Vital Statistics Act could potentially remove millions of names now available for research online. Notwithstanding the importance of these bills to the genealogical community, the House Ways and Means Committee has refused to allow anyone from the genealogical community to testify at the hearing on these bills. As I would be wont to say, the unmitigated gall of these people.
Oh well, here we go again. The price of freedom is eternal vigilance.
Now, on to the other important issues. The U.S. Congress is right now contemplating the passage of two pieces of legislation, one concerning the Social Security Death Index (SSDI) and the other the Model Vital Statistics Act. Passage of either or both of these bills could seriously and very detrimentally affect the entire genealogical community. I will have a lot more about these bills in the immediate future, but right now, information about the bills can be obtained from the House Ways and Means Committee website.
The changes to the SSDI could totally eliminate this valuable resource. The Model Vital Statistics Act could potentially remove millions of names now available for research online. Notwithstanding the importance of these bills to the genealogical community, the House Ways and Means Committee has refused to allow anyone from the genealogical community to testify at the hearing on these bills. As I would be wont to say, the unmitigated gall of these people.
Oh well, here we go again. The price of freedom is eternal vigilance.
Arriving at RootsTech 2012 Wednesday Evening
Yes, in case you were wondering, Salt Lake is still here. South
Temple looks about the same except the construction of the City Creek
Center is winding down. The condominium tower on the corner of South
Temple and West Temple is pretty imposing. The Weather Channel predicted
up to 3" of snow this afternoon and tonight. But, as is usual in Utah
and most other places, they have changed their collective mind and
aren't sure it is going to snow at all. It is still way too warm for
snow.
The onsite registration started for RootsTech and we went in to pick up our portion of the registration. I talked briefly to Rob Holcombe from BYU who is in charge of the event and facilities and he said he thought about 4300 people had registered so far and he expected more tomorrow. I was guessing around 5000 attendees and I just might be correct or pretty close.
One suggestion, if you are coming, don't rely on the elevators in the Salt Palace. They are slower than molasses in January (oops! February). If you need them then get there early and fast.
The onsite registration started for RootsTech and we went in to pick up our portion of the registration. I talked briefly to Rob Holcombe from BYU who is in charge of the event and facilities and he said he thought about 4300 people had registered so far and he expected more tomorrow. I was guessing around 5000 attendees and I just might be correct or pretty close.
One suggestion, if you are coming, don't rely on the elevators in the Salt Palace. They are slower than molasses in January (oops! February). If you need them then get there early and fast.
How many records? A dilemma
Online digitized record repositories' claims regarding their
collections are hopelessly muddled and create a dilemma for genealogists
trying to make a comparison between the online websites. At the root of
the problem is the total lack of consistency about such terms as
record, file, document, names, individuals, collections, and many other
similar terms. Unfortunately, users of the various sites sometimes judge
the relative usefulness of the information based on the way the size of
the database is expressed. This is especially true of family tree or
user contributed sites. Large is equated with good and useful even if
this is not necessarily so.
Let's look at one site for an introductory example.
WeRelate.org is an extremely useful and focused site for displaying individual and family information. The wiki format allows for intensive sourcing and inclusion of media. I highly recommend the site. Now, what about their claims regarding records? Here is a statement from the WeRelate.org startup page, where WeRelate claims to be "the world's largest genealogy wiki with pages for over 2,153,200 people and growing. This is quite an impressive number until you look at the WeRelate.org Special:Statistics page of the wiki. Here is the quote from the page:
This issue of content is not at all unique to WeRelate.org (and I am not picking on that website at all, merely using it as an example). Take for example a comparison between two huge online genealogy giants; FamilySearch.org and Ancestry.com. A superficial look at the two sites would have you believing these two claims: FamilySearch.org claims 1033 "collections" and Ancestry.com claims 30,554 collections. Are the two claims accurate and if they are, do they reflect the relative size differences between the two databases?
It turns out that the term "collection" as used in the two databases are substantially different in their application to the records contained in the databases. Both websites use the term in a totally ambiguous way that gives the user little information about the amount of information on the website.
FamilySearch.org uses the term "collection" in a loose way to designate geographically related records created in a certain way. The term collection is used to refer to original source records as well as extracted records and indexes. So in one instance, the 1855 Alabama State Census is said to contain 34,978 records but in this case, this collection is an index, so it contains names not records. It is certainly not clear what is meant by the term record when each index entry is counted as a record. In another example, the 1869 Argentina National Census is said to contain 1.799,773 records on 157,426 images. Apparently, the number of records refers to the entries on the Census records. But in another collection, such as the Argentina, Salta, Catholic Church Records, 1634 - 1972 there is no number for the records, just a reference 144,293 images. So the total number of "collections" is arbitrary and meaningless. If you drill down into the records, especially those that have images only, you will find that some individual collections are comprised of dozens of rolls of microfilm.
Again, I am not criticizing FamilySearch or anyone else, merely commenting on the vague and ambiguous nature of the designations. Why give a number if the number is meaningless?
Ancestry.com has the same issues as FamilySearch.org but in most cases it is harder to penetrate the confusion. Collections in Ancestry.com are listed with a number of "records." But the number of records is not further defined as pages or individuals or whatever. One number sticks out, the number of records in Member Family Trees is claimed to be 1,838,295,985. Hmm. That is a really big number. How many unique individuals are represented by that number? For example, if I search for one of my ancestors in the Public Member Family Trees, take Henry Tanner for example, I find 57,210 instances of his name. Speculating, if I divide the total number of records claimed by Ancestry.com by the number of duplicate records for Henry Tanner, I get about 32 million entries, still a large number but what is the real number? How many duplicates are there? Isn't this the same problem I started out with on WeRelate.org? Only Ancestry.com does not bother to tell us how many records have content?
The number of records claimed by both FamilySearch.org and Ancestry.com do not give us any idea of how many duplicate records there are for an individual. For example, my ancestor might appear in multiple family trees, but he may also appear in multiple records, all with exactly the same information such as a death certificate and an index of deaths.
The confusion in the terms "record" and "document" is even more dramatic. Fold3.com is an example of using all terms interchangeably. For example, Fold3.com has a list of "collections," claims to have 86,022,535 images, and 100,232,144 memorial pages. Fold3.com collections include American Milestone Documents and Matthew Brady Photographs among other collections. How do we compare the numbers to either Ancestry.com or FamilySearch.org? The simple answer is we can't.
Numbers don't lie, but they don't say much either.
Rather than take these numbers, no matter where they originate, with a grain of salt, perhaps we need a whole salt shaker.
Let's look at one site for an introductory example.
WeRelate.org is an extremely useful and focused site for displaying individual and family information. The wiki format allows for intensive sourcing and inclusion of media. I highly recommend the site. Now, what about their claims regarding records? Here is a statement from the WeRelate.org startup page, where WeRelate claims to be "the world's largest genealogy wiki with pages for over 2,153,200 people and growing. This is quite an impressive number until you look at the WeRelate.org Special:Statistics page of the wiki. Here is the quote from the page:
There are 6,111,311 total pages in the database. This includes "talk" pages, pages about WeRelate, minimal "stub" pages, redirects, and others that probably don't qualify as content pages. Excluding those, there are 2,876 pages that are probably legitimate content pages. (emphasis in the original).What is a "legitimate content page" and how does that differ from the claim to "over 2,153,200 people?" No where is the seeming discrepancy explained. What value are the "people pages" if they contain no content?
This issue of content is not at all unique to WeRelate.org (and I am not picking on that website at all, merely using it as an example). Take for example a comparison between two huge online genealogy giants; FamilySearch.org and Ancestry.com. A superficial look at the two sites would have you believing these two claims: FamilySearch.org claims 1033 "collections" and Ancestry.com claims 30,554 collections. Are the two claims accurate and if they are, do they reflect the relative size differences between the two databases?
It turns out that the term "collection" as used in the two databases are substantially different in their application to the records contained in the databases. Both websites use the term in a totally ambiguous way that gives the user little information about the amount of information on the website.
FamilySearch.org uses the term "collection" in a loose way to designate geographically related records created in a certain way. The term collection is used to refer to original source records as well as extracted records and indexes. So in one instance, the 1855 Alabama State Census is said to contain 34,978 records but in this case, this collection is an index, so it contains names not records. It is certainly not clear what is meant by the term record when each index entry is counted as a record. In another example, the 1869 Argentina National Census is said to contain 1.799,773 records on 157,426 images. Apparently, the number of records refers to the entries on the Census records. But in another collection, such as the Argentina, Salta, Catholic Church Records, 1634 - 1972 there is no number for the records, just a reference 144,293 images. So the total number of "collections" is arbitrary and meaningless. If you drill down into the records, especially those that have images only, you will find that some individual collections are comprised of dozens of rolls of microfilm.
Again, I am not criticizing FamilySearch or anyone else, merely commenting on the vague and ambiguous nature of the designations. Why give a number if the number is meaningless?
Ancestry.com has the same issues as FamilySearch.org but in most cases it is harder to penetrate the confusion. Collections in Ancestry.com are listed with a number of "records." But the number of records is not further defined as pages or individuals or whatever. One number sticks out, the number of records in Member Family Trees is claimed to be 1,838,295,985. Hmm. That is a really big number. How many unique individuals are represented by that number? For example, if I search for one of my ancestors in the Public Member Family Trees, take Henry Tanner for example, I find 57,210 instances of his name. Speculating, if I divide the total number of records claimed by Ancestry.com by the number of duplicate records for Henry Tanner, I get about 32 million entries, still a large number but what is the real number? How many duplicates are there? Isn't this the same problem I started out with on WeRelate.org? Only Ancestry.com does not bother to tell us how many records have content?
The number of records claimed by both FamilySearch.org and Ancestry.com do not give us any idea of how many duplicate records there are for an individual. For example, my ancestor might appear in multiple family trees, but he may also appear in multiple records, all with exactly the same information such as a death certificate and an index of deaths.
The confusion in the terms "record" and "document" is even more dramatic. Fold3.com is an example of using all terms interchangeably. For example, Fold3.com has a list of "collections," claims to have 86,022,535 images, and 100,232,144 memorial pages. Fold3.com collections include American Milestone Documents and Matthew Brady Photographs among other collections. How do we compare the numbers to either Ancestry.com or FamilySearch.org? The simple answer is we can't.
Numbers don't lie, but they don't say much either.
Rather than take these numbers, no matter where they originate, with a grain of salt, perhaps we need a whole salt shaker.
Subscribe to:
Posts (Atom)