Jump to content
Seriously No Politics ×

Two voices in the Book of Mormon


robuchan

Recommended Posts

Posted

...

The cloud (even without Isaiah/Malachi) is so huge compared to those of other authors that I am convinced

it is the work of several authors.

...

I'm in the midst of several activities, so I'll have to get back

to my responses here in a day or so.

Your above comment gives me some hope, though. I've grown old battling

the Brodieites, I fear.

UD

Posted
Chris Smith, on 06 June 2010 - 02:15 PM, said:

If I knew the answer to that, I'd have used it already. I'm not sure there is a perfect methodology for this purpose, but then I'm not an expert in statistical authorship attribution.

Peace,

-Chris

:P Excellent work, Chris. I am on board with you here. Especially with a reticence to accept wordprint studies as very reliable.

I have enjoyed this thread more than any of the previous threads on Book of Mormon authorship, mainly because people with some training in the pertinent fields have weighed in on the subject. I am beginning to think that Craig Criddle's efforts in this area may bear some fruit, in that the methods that his group used will probably be refined and add to the work on word print studies to make them more reliable. After all, the Book of Mormon does invite investigation. The better the critics are, the better will be the responses. I have no problem with a sincere critic.

Glenn

Posted

I'm in the midst of several activities, so I'll have to get back

to my responses here in a day or so.

...

Here is a continuation of my thoughts from yesterday, Bruce:

I

Posted

Hi Bruce,

Excellent post, in messages boards such as this one I find it helpful to have some inclination of what position my interlocutor is coming from so as background info about me with regards to the truth claims of the LDS church I am more or less agnostic. I think the LDS church is a wonderful institution that teaches many good things but whether its teachings accurately describe how the universe really is I have my doubts. I am a theist I do believe in God but I really struggle with the idea of a one true church with exclusive authority from God. I am a Phd student in the sciences I do have some familiarity with machine learning classification techniques hence found the Jockers et al study to be interesting.

It's clear from this plot why you got your results; the Rigdon texts are closer to many of the Book of Mormon chapters than the other authors. But the salient point is that they are not close at all. None of the non-Isaiah BOM texts fall anywhere near the Rigdon cloud.

I to have my doubts concerning the Jockers study but I did have one question concerning your point quoted above. If I understand you correctly you are pointing out Rigdon's cloud of points is indeed closer to the BOMs point cloud hence explaining why he was selected preferentially compared to the other authors. This makes sense as far as it goes but when looking at the Spaulding point cloud it really does not appear all that different from Pratt's and yet Spaulding is selected as a BOM chapter author by NSC far more that Pratt do you have a possible explanation? Now with regards to your overall argument I have a tendency to agree both Spaulding and Rigdons clouds look to far away from the BOM to claim they were authors. But I was puzzled to see the Spaulding point cloud looking not all that different from Patts but Spaulding being selected as an author far more than Pratt. It makes me think some pertinent information is not being captured by the PC analysis.

All the Best,

Uncertain

Posted

Noel00,

I see that you are reading this discussion board.

I would like to ask Craig why he has never (to my knowledge) attempted to answer the main criticism of his work, namely that his authorship attributions are meaningless since his closed-set NSC method is forced to pick a winner out of the candidates he included. Jeff Lindsay made this argument eloquently on his blog (Mormanity), Ben Maguire put forth the same argument on this thread , and other bloggers (whose names I can

Posted

...

If I understand correctly Jockers et al argument is that if none of the tested authors were actually authors of the BOM you would expect a uniform distribution of authorship attribution. That is each author would have a 1/7th probability of being selected as the author

Posted

I keep wondering -- if many, many additional potential authors' word-prints

were added to the study -- whether or not Rigdon and Spalding would still come

out "on top" in so many of the Standford team's BoM chapter attributions?

I also wonder -- if all of those additional word-prints were added, whether

or not there would still be such large statistical "gaps" between those two

authors' "strongly" attributed chapters, and the degree of attribution assigned

to various other potential authors?

I'm not so much impressed with Spalding and Rigdon coming out "on top" with

numerous substantial statistical "gaps," with only a handful of examined

authors, as I would be if the same held true, with many more word-prints added.

I'm curious about the mathematical and statistical meanings of those "gaps."

UD

.

With regards to the gaps my intuition would be this is dependent to some extent on how sensitive the method used is. For example suppose Rigdon is just barely closer to the BOM than the other authors the method is designed to identify which authors from a set of authors is closest to the unknown writing sample. If the method is really sensitive it might pick up on very small differences and magnify them after all its job is to identify differences in writing styles. Hence a large gap may not necessarily mean that Rigdon is much closer to the BOM than the rest it might just reflect the fact that NSC is very good at locating and identifying differences in writing styles even very small differences.

All the Best,

Uncertain

Posted
Now this argument critically depends on assuming a uniform distribution of authorship attribution would be observed if none of the authors tested are actually authors.

I have performed several Delta analyses and found this assumption to be invalid. Some authors' styles are simply more similar than others', and as a result Delta very often chooses one or two clear "winners" out of a limited set of candidates. Here's a very simple example I posted at MDB (copy the link and remove the spaces to make it work):

http://mormon discussions .com/phpBB3/viewtopic.php?p=307972#p307972

Posted

I have performed several Delta analyses and found this assumption to be invalid. Some authors' styles are simply more similar than others', and as a result Delta very often chooses one or two clear "winners" out of a limited set of candidates. Here's a very simple example I posted at MDB (copy the link and remove the spaces to make it work):

http://mormon discussions .com/phpBB3/viewtopic.php?p=307972#p307972

Something is funky with your link.

Posted
Something is funky with your link.

Not if you copy and paste it into the address bar and then remove the spaces. I had to do it this way because the MADB board software blocks direct links to MDB.

Posted

I have performed several Delta analyses and found this assumption to be invalid. Some authors' styles are simply more similar than others', and as a result Delta very often chooses one or two clear "winners" out of a limited set of candidates. Here's a very simple example I posted at MDB (copy the link and remove the spaces to make it work):

http://mormon discussions .com/phpBB3/viewtopic.php?p=307972#p307972

Hi Chris,

Interesting, I can't say I'm surprised as I pointed out I think Jockers et al own work supports this conclusion. Clear winners can emerge even if none of the authors tested actually wrote the text being tested.

I would be curious to see if the BOM internal authorship is consistent. That is do the writings attributed to say Nephi always look like other Nephi writings? Can it be shown that only one writer was solely responsible for writings attributed to Nephi and another different writer was solely responsible for writings attributed to Moroni etc. Does the internal authorship pattern of the BOM follow what the book claims it's authorship patterns should be?

All the best,

Uncertain

Posted

...

it might just reflect the fact that NSC is very good at locating and identifying differences

in writing styles even very small differences.

...

That is rather disconcerting.

In a given case, the data output may show a very low level of

attribution for Barlow and Longfellow --- but the Rigdon data

might be practically identical (say, a 1% true difference).

And yet, when tabulated/charted, Rigdon's degree of attribution

might appear many, many times greater than that assigned to

Barlow or Longfellow.

If this is truly the case, then there is NO statistical difference

of any substance, between the various attributions for a chapter.

And, if that is truly the case, then Bruce's chart (showing such a

wide separation between the BoM chapters' cluster and all the other

authors included) begins to make some sense in a larger context.

How do we explore this possibility?

UD

Posted

That is rather disconcerting.

In a given case, the data output may show a very low level of

attribution for Barlow and Longfellow --- but the Rigdon data

might be practically identical (say, a 1% true difference).

And yet, when tabulated/charted, Rigdon's degree of attribution

might appear many, many times greater than that assigned to

Barlow or Longfellow.

If this is truly the case, then there is NO statistical difference

of any substance, between the various attributions for a chapter.

And, if that is truly the case, then Bruce's chart (showing such a

wide separation between the BoM chapters' cluster and all the other

authors included) begins to make some sense in a larger context.

How do we explore this possibility?

UD

Hi Uncle Dale,

Well it must be remembered my comments were off the cuff. To what extent a large difference in probabilities assigned by the NSC method between the first and second ranked authors represents the "true" difference between the two I have no idea. To me it seems reasonable the difference in probabilities will be at least somewhat dependent on how good the method is at identifying differences in writing styles if the method is really good a large difference in probability may be just a small difference in writing style. But this is not necessarily bad maybe there is a small but relevant difference in writing styles and the method is just very good at picking up small relevant differences. Or maybe the method is very poor at picking up differences in writing styles in which case a large difference in assigned probabilities between the authors represents a very large difference in writing style. To make a long story short I am not comfortable at this time stating a large different in assigned probabilities represents a large difference in writing styles. But heck maybe an expert in NSC like for example Daniela Witten :P might be so comfortable.

(Edit: I would also add what is important is to what extent large differences in assigned probabilities correctly represent who really wrote the given text. If large differences in probabilities only represent very small differences in writing style if the method correctly picks the right author this is what is important)

All the Best,

Uncertain

Posted

...

maybe an expert in NSC like for example Daniela Witten :P might be so comfortable.

...

I guess we can always hope for her avatar's benefic appearance here.

(I'm rather good at hoping.)

UD

Posted

Craig sent me a response to Bruce, however since I was reading it on webmail at work during lunch, the formating needed correcting. Maybe when I use my home PC it will look better and I'll post it in about 5 hours. (after I watch the Daily Show. :P

Posted

Craig sent me a response to Bruce, however since I was reading it on webmail at work during lunch, the formating needed correcting. Maybe when I use my home PC it will look better and I'll post it in about 5 hours. (after I watch the Daily Show. :P

Thanks again for your efforts to make this a very interesting thread, noel00.

Posted

Noel, Here is my response to Bruce. Craig

Hello Bruce,

I'm surprised that you perceive use of a closed set of candidate authors as the "main criticism " of our work. I'm even more surprised that you accuse me of avoiding a discussion of that point. I actually never thought it was a point worthy of debate.

Yes, of course, we tested a closed set of authors. That's true of most authorship attribution studies. It's also true that if our list of authors did not include any of the true authors in the case of multiple authorship or if it did not include the true author in the case of single authorship then the results would not be meaningful . So historical evidence becomes important for selection of candidate authors and for drawing meaningful conclusions.

If you assume a naturalistic view of the origins of the Book of Mormon, as we did, you are left with three theories for the Book of Mormon:

Theory 1. Smith composed it by himself.

Theory 2. Smith had help.

Theory 3. Someone else wrote it. Smith was just a front man.

At the time of the LLC study, we did not have reliable text for Joseph Smith. So we chose Spalding, Rigdon, Cowdery, and Pratt as candidate authors based on historical evidence implicating them. We also included positive controls (Isaiah/Malachi) and negative controls (Barlow and Longfellow). We then ran the tests using Delta (a commonly used method) and Nearest Shrunken Centroids (NSC, a method developed for classification of gene expression patterns) to determine which chapters of the Book of Mormon had frequent word usage patterns --"clouds", as you call them-- that clustered closest to the clouds of our test and control authors. The two methods gave basically similar patterns of attribution, though NSC was more accurate (9% error rate) than Delta (11% error rate) in cross validation tests.

Our study had the potential to disprove the Spalding-Rigdon Theory - say, for example, if the clouds for Barlow or Longfellow were frequently identified as closest to clouds for a large number of chapters of the Book of Mormon, or, for that matter, if the clouds for Cowdery or Pratt were identified as the clouds closest to those of a large number of chapters in the Book of Mormon. But that did not happen. Instead, we found that there were many Book of Mormon chapter clouds that clustered close to the Spalding cloud and many that clustered close to the Rigdon cloud. These attributions were consistent with the Spalding-Rigdon theory. We also had chapter clouds that were close to the clouds for Isaiah/Malachi, including a fair number of false positives for Isaiah/Malachi. The false positives may indicate an effort to imitate the style of the Bible. if so, that would be consistent with historical evidence. Eye witnesses reported that Spalding modified his writing style in Manuscript Found to imitate the "old style" of the Bible, including frequent use of the phrase "came to pass". And, as you have noted (thank you), Rigdon's 1864 revelations, written in his unedited "prophet mode", fall closer to the Book of Mormon than his earlier work that was not written in a scriptural style.

In the time since our LLC publication, my colleague Matt Jockers has identified a body of text that may be Smith's, and he has carried out additional analyses with and without the putative Smith signal. He also omitted isaiah/Malachi as candidate authors and focused on the non-Biblical chapters. This has provided us with a better picture of plausible 19th century contributors -- again under the assumption that the true authors are among those included in our testing. The new attribution pattern turns out to be largely consistent with the earlier pattern, but we now see about 20 chapter clouds that are closest to the Smith cloud. Several of the Smith-attributed chapters make sense in terms of the chapter content and historical evidence. So our current textual evidence is more supportive of Theory 2 than Theory 3, and appears inconsistent with Theory 1. I can also add that there is also new historical evidence that likewise appears more consistent with Theory 2 than Theories 1 or 3.

With respect to your statistical questions, my colleagues have asked that they be addressed through normal professional venues, not on-line forums, such as this.

In the interim, please see preprint pdf files (listed below) for articles that are available at Matt Jocker's web site. All of these articles are accepted for publication in the Journal of Literary and Linguistic Computing. The second article deals specifically with statistical issues.

1. A preprint version of our published LLC publication:

http://www.stanford.edu/~mjockers/pubs/LLCPreprintReassess.pdf

2. A comparative study of machine learning methods for authorship attribution written by Matthew Jockers and Daniela Witten.

http://www.stanford.edu/~mjockers/pubs/LLCPrePrintFederalist.pdf

[Note: in the above study, Matt and Daniela found that nearest shrunken centroids (NSC), one of the methods we used in the LLC paper, had superior performance among the 5 methods evaluated - Delta, k-nearest neighbors (KNN), the support vector machine (SVM), nearest shrunken centroids, and regularized discriminant analysis (RDA). Matt and Daniela also provide a discussion of Principal Components Analysis on pages 12-13].

3. A study of Smith's personal writings using NSC Classification written by Matthew Jockers.

http://www.stanford.edu/~mjockers/pubs/SmithNSCAnalysis.pdf

[Note: in the above study, Matt found that 15 of 96 documents attributed to Smith clustered with text written in Smith's own hand. These and the hand-written documents are the collection of documents that he recently used as a putative Smith signal for re-testing of the Book of Mormon].

With regard to peer review, I'm surprised that you dismiss its value. Of course things do pass through peer review that should not, but, over time, such studies tend to be challenged and overturned. I personally much prefer peer review to the alternative, as it provides a measure of quality control. Without it, where would science be?

Craig

Posted

I'm surprised that you perceive use of a closed set of candidate authors as the "main criticism " of our work. I'm even more surprised that you accuse me of avoiding a discussion of that point. I actually never thought it was a point worthy of debate.

Yes, of course, we tested a closed set of authors. That's true of most authorship attribution studies. It's also true that if our list of authors did not include any of the true authors in the case of multiple authorship or if it did not include the true author in the case of single authorship then the results would not be meaningful . So historical evidence becomes important for selection of candidate authors and for drawing meaningful conclusions.

Hi Craig,

Thanks for taking time to respond, and thanks for addressing the main criticism of your work that has appeared in several reviews of your work over the past year. It is not the only serious criticism of your paper in my view, but it is a very serious flaw.

You are exactly right in conceding (thank you) that

post-16623-127600920246_thumb.png

Posted

I think that it ought to be quite clear to anyone who looks at the charts that Longfellow and Barlow (choices that were made seemingly quite subjectively) were clearly bad choices for a negative controls. Their inclusion couldn't possibly tell us very much - other than to lend a seeming air of accuracy to the study. And I thought that the original study was (as Bruce points out) far too dismissive of the incorrect attribution to the artificial Malachi/Isaiah author. It seems to me that in particular, in the way that vocabulary was selected, negative control authors extend far too much influence on the test in general, and in a way that doesn't return any value for what is costs to the study.

It occurs to me to rephrase something that Bruce wrote earlier. If we use the data plot provided, and we knew the author of the Book of Mormon text, but not the author of the Spalding text, which author would be assessed as most like the unknown Manuscript Found? Would it be the Book of Mormon author? The Longfellow Author? Pratt? Barlow? Cowdery? Rigdon? Isaiah/Malachi?

If we redid this experiment except that we used Rigdon as the unknown author? We certainly wouldn't see a lot of the Book of Mormon in these authors.

Ben M.

Posted

...

Would it be the Book of Mormon author?

...

Somehow that doesn't make any sense to me.

It's like measuring the author (voice) of this

thread against the voice of Ben McGuire, as

found in other web-documents -- to determine

if the author of this thread wrote the other

McGuire texts.

I suppose that the experiment might be conducted

once -- just to prove that Ben is not the author

of all the contributions to this thread -- but after

that initial disproving, any other application would

be meaningless/useless.

In short -- this thread does not have a voice

attributable to a single author, and neither does

the Book of Mormon.

UD

Posted

You are exactly right in conceding (thank you) that

Posted

Dale writes:

In short -- this thread does not have a voice attributable to a single author, and neither does the Book of Mormon.
Right, but this kind of look at the data is author independent. We are simply plotting data points. (We do associate the data points with an author, but, it doesn't matter really in this question whether or not we treat the Book of Mormon as a single voice or as multiple voices). Now, given that, how many of the Book of Mormon plots would find themselves preferred with respect to an unknown author of the Spaulding manuscript compared to say the data plots from Barlow or Longfellow? That is the issue of the clustering. With one or two exceptions (in the Isaiah sections - see how the Book of Mormon data points overlap the Isaiah data points in those places), Spaulding is almost always closer to these other authors than it is to the Book of Mormon. Not just a little closer - much, much closer. What do you think would be a reasonable basis to exclude an author? Suppose that we take just the plots of the Spaulding data. We find an approximate center. We calculate the distance to its most distant point. Now, how much farther do you think we could go before we conclude that another data point is clearly not likely to be a part of that same authorship? Twice the distance? (That is, take the furthest point from the middle of the Spaulding cluster and go that far again?). At what point can we conclude that even if the NCS predicts authorship, that there is no likelihood of authorship, and so we can only conclude that the author is not in the closed set?

Ben McGuire

Posted

...

At what point can we conclude that even if the NCS predicts authorship, that there is no

likelihood of authorship, and so we can only conclude that the author is not in the closed set?

...

Again, we are running up against this term, "the author."

If you write a text, from beginning to end, then we might call

you "the author." But if you compose a long text including

quotes and paraphrases from Helaman and General Moroni, and

that narrative is later edited by Mormon, then we are back

to the same problem I was talking about in my earlier reply.

I do not think that the Delta or NCS authorship attributions

can be taken as "proof" of anything. The attributions supplied

by those methods do, however, form patterns of distribution across

the Book of Mormon text --- something is being measured, and

showing a rough view of that book's components' literary "texture."

With this recent pca plotting we are measuring something else,

but again we can obtain another rough view of that "texture."

I am not surprised to see differing results, when different things

are being measured.

If you read back in this thread, you'll see that I recommended that

Bruce transfer his x,y measurements to a bar graph incorporating

each BoM chapter in order -- with the x component depicted as color

and the y component depicted as bar height. Once that graphic is

available, we will have at our disposal yet another rough view of

the "texture" patterns discernible across the Book of Mormon.

There is a method inherent in what I am suggesting here.

Back when I was a grad student in the Geography Dept. of the

UofU I was involved peripherally in a NASA project involving

the correlation of satellite remote sensing data of the Great

Salt Lake area. We had on hand imagery from such sources as

radar, infra-red, magnetic, microwave, visible light, etc. etc.

Our initial task, as I recall, was to "scale" the imagery so

that the various data sets could be superimposed, one atop

another. In some instances there were no discernible features

in common between data sets. This was back in the days before

gps coordinate positioning, and it was hard work getting the

several sets of images coordinated for superimposition.

When at last, a sample quad of multiple layers of different

imagery was produced, it was quite exciting to see the results.

Features not discernible in one set of images were discernible

in other kinds of image -- water salinity, for example -- or

water current patterns -- or brine ship habitat locations.

I suggest that we apply a similar set of examinations to the

Book of Mormon chapters (or better yet to the 1830 ed pages),

and build up a set of different sorts of views of that text.

In such a project we should not expect one view to reveal the

exact same literary "texture" as another view. But in the end,

some combination of overlays of quantified examinations should

tell us things about the literary composition that we cannot

see simply by reading the narrative and making subjective

comparisons with the writings of other purported authors/sources.

From that perspective, I'm overjoyed to be able to look at

Bruce's chart. I want to see more pca charting, based upon

different sorts of data comparison relating to the BoM.

But, you'll probably understand why I cannot accept that same

isolated chart as constituting preemptive proof that the

Book of Mormon could not contain textual contributions from

Spalding, Rigdon, Smith, etc.

UD

Posted
We calculate the distance to its most distant point. Now, how much farther do you think we could go before we conclude that another data point is clearly not likely to be a part of that same authorship? Twice the distance?
Part of the problem with amateur statisticians is that they tend to look at "clouds" of data-points, can make no sense of such charts, and conclude that there is no relationship between the variables thus charted. However, statistics can mathematically analyze such "clouds," and say that there probably is, or is not, a significant relationship. As an example, the numbers representing the overlap between Dale's and Jockers' identification of Spalding-esque units of the book may not look impressive, but when one statistically analyzes them, the results are impressive. We look at groups of data-points, not individual ones, unless they are extreme outliers.

Your comments are appreciated, Ben.

Archived

This topic is now archived and is closed to further replies.

  • Recently Browsing   0 members

    • No registered users viewing this page.
×
×
  • Create New...