Jump to content
Seriously No Politics ×

Another Wordprint-type Study


LifeOnaPlate

Recommended Posts

Posted
I am curious to know what his response to your assessment has been.

I'm not asking for any privileged information here -- just a general statement

on whether or not the Stanford researchers agree that you are raising some

valid points, worthy of consideration, in conducting any similar, new studies?

UD

Hi Uncle Dale,

To be honest I have no idea :P. To be frank I am not an expert in word print studies or a statistician. I do have some experience working with machine learning/classification techniques. And I have been trained to some extent in statistics. But I do not claim expert status in either classification or statistics. My assessment is based on my understanding of the paper plus my understanding of the nearest shrunken centroid algorithm. I am certain a true expert could no doubt find flaws in my reasoning and find additional concerns about the manuscript I have not seen. No study is going to be without flaw if you look hard enough there is always ways to improve. I personally do not think this study is a good reason for LDS members to en mass turn in their membership. But given my pluristic tendencies I don't think conclusively proving the BOM as a fraud would necessarily be a good reason to stop being LDS. I think this manuscript is a good debate starter and I hope it leads to similar peer-reviewed studies published in mainstream scholarly journals both pro and con.

All the Best,

Uncertain

Posted
Hi Selek,

I do not agree. If there is not enough data to estimate Joseph Smiths word print there is not enough data. If there is not enough data including him would be pointless any probabilities calculated for JS as author would have errors so large the estimated probability would be worthless.

I am not sure you are following my argument. If I understand the paper correctly the authors are claiming each putative author is assigned an absolute confidence value (i.e. probability) which represents how confidently the unknown passage is assigned to the given author. For your example above yes Thurber may very well be the author that is calculated as the most probable among the set of authors tested. But if the method is working correctly the calculated confidence value would very small. That is Thurber would be one poor choice out of many poor choices. So yes the results would have a connection with reality. They would indicate Thurber was the best choice out of the available options but that given the low confidence value Thurber is unlikely to have written the tested book. This in no way refutes Jockers et al conclusions. Now if you are asserting the method is producing flawed confidence values I am listening but this must be demonstrated using solid science (like Ben is trying to do) not just asserted.

In addition suppose JS was included I dare say you would be arguing the "real" authors (i.e. Nephi, Alma etc. ) were not included therefore we should disregard this study.

All the Best,

Uncertain

You're probably right- but this study as constituted has no more reasonable bearing on who actually wrote the Book of Mormon than had they simply grabbed three random passers-by off the street and asked them to submit writing samples for analysis.

The numbers they are arriving at are dictated by their sample. If the sample is reasonably complete, the results are reliable. If the sample is deliberately skewed or incomplete, then the numbers are meaningless.

This is the same phenomenon we see with the political polling during the last election- skewed samples invariably return skewed results- no matter how diligent the analysis.

You're attempting to hide from a binary solution set. Either the real author(s) are included in the sample or they are not.

If the "real" authors aren't included in the sample, then the numbers themselves on "probable authorship" are meaningless.

The authors of this study are proceeding from a priori assumptions and then assigning meaningless statistical numbers and pretending that they've revealed something.

If you start with the assumption that the Earth is the center of the universe and reject all data which contradicts your hypothesis, then your conclusions have no meaning- no matter how carefully you measure the heavens.

The simple facts of the matter contradict the Spaulding hypothesis and the contributions of the other authors except Rigdon, as has been demonstrated many times.

As has also been pointed out several times in this thread, this study is little more than an attempt to make new data fit an old (and discredited) theory.

Posted
You're probably right- but this study as constituted has no more reasonable bearing on who actually wrote the Book of Mormon than had they simply grabbed three random passers-by off the street and asked them to submit writing samples for analysis.

Their numbers they are arriving at are being derived from their sample. If the sample is skewed, then the numbers are meaningless.

You're attempting to hide from a binary solution set. Either the real author(s) are included in the sample or they are not.

If the "real" authors aren't included in the sample, then the numbers themselves on "probable authorship" are meaningless.

The authors of this study are proceeding from a priori assumptions and then assigning meaningless statistical numbers and pretending that they've revealed something.

The simple facts of the matter contradict the Spaulding hypothesis and the contributions of the other authors except Rigdon, as has been demonstrated many times.

As has also been pointed out several times in this thread, this study is little more than an attempt to make new data fit an old (and discredited) theory.

Hi Selek,

Your argument appears to be that without including the "real" author any confidence values calculated are worthless. This is contrary to what the authors are apparently claiming in the manuscript. As I already stated I am open to your argument but so far all I have seen is assertion (i.e. JS must be included in order for the confidence values to be accurate). I need something more than that. Based on my understanding of the NSCC algorithm confidence values are based mostly on how close a given unknown text matches the centroid. If the unknown text belongs to none of the possible authors it will not match closely with any of them. Hence the confidence values calculated for all authors will be very small it is irrelevant whether or not the "true" author is present. If you are interested it is equations 6, 7, an 8 in this manuscript:

http://projecteuclid.org/DPubS?service=UI&...d.ss/1056397488

That outline how the relevant probabilities are calculated.

All the Best,

Uncertain

Posted
...

Hopefully you haven't abandoned work on the biography for this lesser task.

...

Not abandoned -- but greatly postponed, mostly due to health problems

(my own and other family members').

Basically I'm stuck in 1822, trying to determine who Rigdon was associating

with in the Redstone Baptist Association, in and around Pittsburgh.

Unfortunately, I'm no longer able to travel to Pennsylvania to consult the

necessary manuscript materials there.

Thus, I waste my time on web MBs, etc., instead of contemplating the obscure

activities of one, Elder Sidney Rigdon.

Oh well, maybe Kevin will convert to RLDS and offer to do my library research

in Pittsburgh for me ---- one can always hope.

UD

Posted

Uncertain writes:

I do not think excluding Joseph Smith (JS) as a potential author is a serious flaw. I thought the explanation given by Jockers et al was a good one. Apparently they could not find enough material certain to have been written by Joseph Smith hence they did not include him. This seems reasonable to me. If as the authors claim they are calculating the real probability of a particular author being the unknown author of a passage. Then including additional authors should in theory make no real difference. In other words if a passage has a 80 percent probability of being written by Sidney Ridgon (SR) than regardless of how many other authors are included the 80 percent probability should not change much. For that matter if SR was the only author being compared the absolute probability should still be around 80 percent.
The issue remains though, that despite their claim that they can produce a fixed probability, they method they use doesn't actually produce a fixed probability. It can only produce a relative probability.
So we are agreed contrary to earlier reports the authors do claim to be calculating the real probability not the relative probability.
No, in fact, I am stressing the point that this study cannot produce a real probability - only a relative probability. And of course this is why excluding Joseph Smith is such a critical flaw. I don't think that there is any solid basis on which to limit the possible authors of the Book of Mormon to this short list. I believe that the majority view (notwithstanding the voices on this forum) among LDS and their critics is that these men were NOT the authors of the text (see Vogel and others for the critical perspective). And so for this argument to hold much meaning, you would first need to demonstrate (i.e. provide real evidence) that these men were the authors so that then we could take the text apart in this way and determine which sections were authored by which individual. But, in terms of trying to prove (as evidence) that these men were the authors, this study has absolutely no merit.
Absolutely I agree with this reasoning I make the same point above. If indeed the method is producing a real probability it should give the same result for known author B regardless of how many authors are included in the analysis. This is the reason as I pointed out I do not view excluding JS as a serious flaw.
The challenge, of course, is that this study won't produce the same results if we modify the list of authors, the sample texts, and the controls.
Personally I have no problems using a cutoff of .1%. NSCC scales each element of each centroid by the standard deviation of the element for all samples. So for example if the frequency of a given word X varies wildly in different samples of the same authors work. The frequency of occurrence of this word would have a large standard deviation and hence the importance of this word in classifying unknown samples of writing would be greatly diminished (i.e. this word would have little impact on deciding how closely an unknown sample is to the given authors work). In order to calculate such parameters as the standard deviation you need enough samples hence the .1% cutoff I donâ??t see the problem.
The issue is that the other samples - including the control samples - will influence which words in the vocabulary are kept and which are discarded. This means, of course, that using a different set of samples - particularly using a different set of control samples - will result in a completely different set of variables and hence a completely different experiment. This is not, ideally, what a control is supposed to do.

A good control for this experiment might be to eliminate the most common elements - that is, instead of using an actual text and author as a control, evaluate a composite of hundreds of authors to get a relative frequency list (the most frequently used 200 words in 1830s english for example), and then creat a faux author that uses this model. That way, we become interested more specifically in those features of a particular sample which deviate from what we might call the average language.

I do not agree I think using random controls would be a bad idea. The ideal control is the control that is identical to your test case except for differing in the variable you are trying to test. So if I am testing a new drug I want my control cases to be identical to my test cases except for they are given a placebo. The authors tried to select controls as similar as possible to the BOM indeed they performed hierarchal classification showing the control texts were more similar to the BOM than 50 other novels from the same era. This is entirely appropriate you want your controls to be as close as possible to your test cases. I see no justification for accusing them of swapping out controls until they got the required results. This would constitute scientific misconduct and I am not going to make this accusation without rock solid evidence.
The problem, of course, is that according to you, it shouldn't matter which control texts I use because, of course, the probabilities are supposed to be consistent. This means that a random sampling should do just fine, right? In reality, a random sampling here significantly modifies the experiment. And this suggests that the results aren't real probabilities but only relative probabilities. Do you see the point? The bigger problem is that the controls they used weren't selected on the kind of basis you suggest, and in fact we are given some reasons why they were used, but those reasons don't seem to have much to do with the study itself. Choosing a poetic work because its styled poetically and has some repitition doesn't seem to be good criteria, nor does selecting a text which uses the word "and" frequently to start phrases - particularly in a test which is all abotu vocabulary and frequency. Wouldn't such a choice dilute the value of the word "and" (which was in fact one of the words in the tested vocabulary).
Which author ranks first second third etc. is certainly relative it depends on which authors are including in the study. But the probabilities should not be nearly as relative. If for example an unknown tumor sample is analyzed but its corresponding cancer is not included in the study. The tumor will indeed be assigned to an incorrect cancer but its probability of being of that cancer type should be very low (if the method is working properly).
But NCSS doesn't give absolute probabilities in a case like this. It tells us in a vector type calculation which one is closer to the unknown - and gives us a general mechanism to compare to other classes in detrermining relative "distances" (although using the notion of "distance" is somewhat misleading). In fact, it tells us these things even if we exclude the known author from a test case. We could randomly pick 15 authors, draw up a sample, and still get the kinds of numbers this case produced. This type of experiment ONLY works if we are already certain that the class is among those being sampled.
It might be a wise idea to run your methodology with the seven authors used in the study on the BOM to make sure you duplicate the studyâ??s results. You can then add additional authors and see how things change.
This is more of a challenge without getting their data, since they don't really give detailed information in the study about the exact contents of the samples. I would need their matrix (presumably in XML format) to duplicate that portion of the task. Myself, I have a number of lexical tools, and can easily produce what I need from the sources I have (my programming is done in this case as forms in a Access database, giving me easy exports into Excel for analysis).
Now it might be the authors did include probabilities in their attribution of authorship but if so I missed it when I read the paper.
And the reason why this wasn't there is because their method doesn't actually calculate real probabilities ... it simply isn't there.

From the paper, under results, we get this admission:

For each chapter of the Book of Mormon, using both NSC and delta, we compared the relative probability that a candidate author or a control author contributed to that chapter.
And there it is. The relative probabilty (somehow I missed that again this morning - I knew it was there someplace).

When they talk about errors, they suggest that:

The lowest NSC error rate was obtained when all 110 words were included; the error rate was 8.8%. This means that we would expect to classify correctly a new sample written by one of the seven candidate authors 91.2% of the time. Since there are seven candidate authors, a classifier that selected an author completely at random would give a correct classification rate of 1/7 or 14.3%, and an average misclassification error rate of 6/7, or 85.7%. Therefore, the low error rates obtained using NSC and delta are impressive.
Here is the challenge. The probability of them picking the correct author if the author is not in the sample is identical to the odds of getting the correct author by choosing one completely at random: 0%.

This doesn't mean that the test wouldn't be without value. This kind of approach could be very helpful in analyzing, say, the Federalist Papers, whose disputed authorship can be assessed against a fairly tight list of possible authors. But here, unless we can somehow demonstrate that Rigdon et al. were the only options for authorship, this study doesn't help us confirm that authorship - it can only tell us that if that model of authorship is right, which portions of the Book of Mormon likely came from which sources. And that is all it can do. And in the end, since most of the people in this forum reject the Spaulding/Rigdon authorship model (LDS members and critics), most of the participants here will also reject the conclusions of this study.

Ben M.

Posted
I do not think excluding Joseph Smith (JS) as a potential author is a serious flaw. I thought the explanation given by Jockers et al was a good one. Apparently they could not find enough material certain to have been written by Joseph Smith hence they did not include him. This seems reasonable to me.

Maybe they can use their wordprint method to identify things JS actually wrote (I'm thinking of unattributed editorials in T&S, etc.) :P

Posted
Uncertain writes: So we are agreed contrary to earlier reports the authors do claim to be calculating the real probability not the relative probability.

No, in fact, I am stressing the point that this study cannot produce a real probability - only a relative probability. And of course this is why excluding Joseph Smith is such a critical flaw. I don't think that there is any solid basis on which to limit the possible authors of the Book of Mormon to this short list. I believe that the majority view (notwithstanding the voices on this forum) among LDS and their critics is that these men were NOT the authors of the text (see Vogel and others for the critical perspective). And so for this argument to hold much meaning, you would first need to demonstrate (i.e. provide real evidence) that these men were the authors so that then we could take the text apart in this way and determine which sections were authored by which individual. But, in terms of trying to prove (as evidence) that these men were the authors, this study has absolutely no merit.

Hi Ben,

As always a pleasure to dialog with you.

I am somewhat confused about what your position is. In post #136 you said this:

â??The study actually seems to claim to be calculating the real probability that a particular author is the unknown author of a passage. However, this isn't what they are doing. There are several ways to show that this is the case. This might be the case if they were certain to have included all possible authors in the list of known authors (which is part of the reason for the otherwise non-essential historical narrative).â?

This is the basis for my comment you are responding to. It seemed to me you were agreeing that the authors claimed in the study to be producing real not relative probability. But that you yourself disagreed with them.

In this post you quoted this from the paper:

â??For each chapter of the Book of Mormon, using both NSC and delta, we compared the relative probability that a candidate author or a control author contributed to that chapter.

And there it is.â?

Which makes me think you now believe the authors never claimed to be calculating anything other than the relative probability. Or maybe I just misunderstood your earlier post #136.

With regards to the above quote from the paper I interpreted this as the authors taking the probabilities for each author and comparing them to the control authors producing a relative difference in probability. In other words the control authorâ??s probabilities are very different from the rest of the authorâ??s probability. If all authors are equally not connected with the text this is not the pattern that would be expected. I wager this is why you think they preselected their controls on the basis of the desired results. Something I still think would constitute scientific fraud.

There are many cases in which the authors seem to be claiming a real probability not a relative probability. Or at least they do not qualify their statements such that they refer to it as a relative probability.

Examples:

â??Our work employs two techniques to determine the probability that each chapter of the Book of Mormon was authored by each of seven authors: Oliver Cowdery, Parley Pratt, Sidney Rigdon, Solomon Spalding, Isaiah-Malachi (from the Bible), Henry Wadsworth Longfellow, and Joel Barlow.

â??

â??On this basis, NSC assigns a probability that each potential author wrote each Book of Mormon chapter;â?

But I do see how the wording can be ambiguous. Perhaps we just need to seek clarification from the authors.

The challenge, of course, is that this study won't produce the same results if we modify the list of authors, the sample texts, and the controls.

This is the key issue. If you can demonstrate that modifying the list of authors, sample texts and the controls does produce dramatically different results. This would be very problematic with regards to the claims made in this paper. But I think the analysis needs to be done before we jump to any conclusions.

The issue is that the other samples - including the control samples - will influence which words in the vocabulary are kept and which are discarded. This means, of course, that using a different set of samples - particularly using a different set of control samples - will result in a completely different set of variables and hence a completely different experiment. This is not, ideally, what a control is supposed to do.

Yes but the words are kept on the basis of having enough of them such that meaningful statistics could be done. Sure by adding or removing controls different words may be kept. But is this going to dramatically change the overall results? What I am trying to say is suppose you added a different control this would certainly change which words are kept and which are removed. But all that is important is that enough words are kept to build a meaningful word print for each author. As long as that is achieved which words are used may not matter. What would be significant is if the added control suddenly started being ranked at the top of the author list. But again to build a convincing case the analysis needs to be done.

The problem, of course, is that according to you, it shouldn't matter which control texts I use because, of course, the probabilities are supposed to be consistent. This means that a random sampling should do just fine, right? In reality, a random sampling here significantly modifies the experiment. And this suggests that the results aren't real probabilities but only relative probabilities. Do you see the point?

Yes if indeed the probabilities are real and not relative it should not make a dramatic difference what is used. That is no matter what controls are used they should always be assigned a low probability and the â??realâ? authors should have a high probability. By trying to select as controls texts that are similar to the BOM the studyâ??s authors are simply covering their bases. In other words if they randomly selected a text as a control that is very different from the BOM. Critics could point to it and claim it is not surprising the given text scored so low look how different it is from the BOM. You may not like how the control was selected but I do not think this is reason to start believing they swapped out controls until they got the desired results. I just have a hard time believing Craigs coauthors (and Craig for that matter) would do such a thing. I mean it really is an outrageous maneuver and one easy enough to find out. That I doubt the authors believed they could get away with such a thing.

But NCSS doesn't give absolute probabilities in a case like this. It tells us in a vector type calculation which one is closer to the unknown - and gives us a general mechanism to compare to other classes in detrermining relative "distances" (although using the notion of "distance" is somewhat misleading). In fact, it tells us these things even if we exclude the known author from a test case. We could randomly pick 15 authors, draw up a sample, and still get the kinds of numbers this case produced. This type of experiment ONLY works if we are already certain that the class is among those being sampled.

I am not convinced this is the case. I took another look at how probability is calculated using NSCC. I wish this board had an equation editor so I could show it on the board. But in broad terms the key equation (equation 6) consists of two terms the first term measures how different the unknown sample is from the centroid. The second term is the class prior probability which is the overall proportion of class k in the population. The second term gives more weight to those classes with more known examples.

The bottom line is term one depends mostly on how different an unknown sample is from a given centroid. And will probably vary little with the addition of more authors since the difference between an unknown sample and centroid k (i.e. author k) is not affected by the difference between and unknown sample and centroid y. Now the second term does depend to some extent on the other classes which means it is relative. It seems to me how relative a given probability estimate is depends on which term dominates for a given probability calculation.

This is more of a challenge without getting their data, since they don't really give detailed information in the study about the exact contents of the samples. I would need their matrix (presumably in XML format) to duplicate that portion of the task. Myself, I have a number of lexical tools, and can easily produce what I need from the sources I have (my programming is done in this case as forms in a Access database, giving me easy exports into Excel for analysis).

The issue is that the more your method differs from theirs the more the authors can claim an â??applesâ? to â??orangesâ? comparison. This is why it is best to test a methodology as close as possible to the ones used in the study.

When they talk about errors, they suggest that:

â??The lowest NSC error rate was obtained when all 110 words were included; the error rate was 8.8%. This means that we would expect to classify correctly a new sample written by one of the seven candidate authors 91.2% of the time. Since there are seven candidate authors, a classifier that selected an author completely at random would give a correct classification rate of 1/7 or 14.3%, and an average misclassification error rate of 6/7, or 85.7%. Therefore, the low error rates obtained using NSC and delta are impressive.â?

Here is the challenge. The probability of them picking the correct author if the author is not in the sample is identical to the odds of getting the correct author by choosing one completely at random: 0%.

In this case they are talking about error rate among the known samples. So a classification rate of 1/7 for a random classifier is accurate. In other words they tested how accurate their method was when discriminating among a set of known writing samples from seven authors.

To sum up where I am right now I am taking a wait and see attitude. It appears to me the a key point of contention is whether or not real versus a relative probability is being calculated. My reading of the manuscript indicates the authors are claiming a real probability. But the text is somewhat ambiguous and I see how it can be read differently. I took a look at the relevant equations by which the probability is calculated and it also seems ambiguous to me. Therefore I will wait for the verdict of those more versed in this issue than myself. Or for the result of empirical studies demonstrating this method produces dramatically different results based on different controls or added authors. I just do not think at this point there is sufficient evidence to summarily dismiss the study because JS was not included and go home and sleep the sleep of the contented :P.

All the Best,

Uncertain

Posted

Hi All,

So I took yet another look at the NSCC paper and I now believe the probabilities are mostly relative. That is a given probability is relative to all other classes being examined. I simply missed it when I looked at it before. According to how the NSCC paper calculates the relevant probability it looks to me like if there was only one class being compared (i.e. one author). That author will receive a probablity of one regardless of how similar the unknown sample is to the author. Steven King could be compared to the Bible and the probablity calculated would be one. At least according to my interpretation of the relevant equations. I think the only significance to this study is demonstrating that the writing styles of authors like SR etc. are closer to the BOM than a set of control texts. Which is not what you would expect if all authors examined had no connection with the BOM. I am not sure this finding is all that significant. Now it is possible I have misinterpreted the relevant equations and am open to correction from those more versed in this field than myself.

All the Best,

Uncertain

Posted

Uncertain,

I am glad we are both on the same page now.

The value in this study (and I am sure the basis for its being submitted) is that it demonstrates that using this kind of model, we can get a more accurate result than previous kinds of word studies - under the circumstances that we have some idea who the author(s) might be. And I assume that the reason for the historical narrative in the paper was to try and establish that fact for the five authors who were being proposed. And there, I think, is where most people will start to disagree with the conclusions. I think that control texts are simply pointless - and that rather than dealing with the biblical literature by including an Isaiah/Malachi author they should simply have excluded those texts as samples, and they should have probably done an additional study with Joseph Smith using the limited information available (even if it was a small sample size) - just as they recalculated without the 2 controls in the five-author study.

Ben M.

Posted
Uncertain,

I am glad we are both on the same page now.

The value in this study (and I am sure the basis for its being submitted) is that it demonstrates that using this kind of model, we can get a more accurate result than previous kinds of word studies - under the circumstances that we have some idea who the author(s) might be. And I assume that the reason for the historical narrative in the paper was to try and establish that fact for the five authors who were being proposed. And there, I think, is where most people will start to disagree with the conclusions. I think that control texts are simply pointless - and that rather than dealing with the biblical literature by including an Isaiah/Malachi author they should simply have excluded those texts as samples, and they should have probably done an additional study with Joseph Smith using the limited information available (even if it was a small sample size) - just as they recalculated without the 2 controls in the five-author study.

Ben M.

Hi Ben,

Yep I think we are more or less on the same page. This is the second time on this board I have had a long involved discussion and capitulated in the end. Clearly I am no good at this debating business :P. I am not sure I entirely agree that the controls are worthless I think a determined critic could make a limited amount of hay out of the fact that the controls always show up at the bottom of authorship rankings. I am not sure including JS would have proven any more convincing. If he showed up with a high probability I think the believing response would simply be the "real" authors are not included in the analysis.

All the Best,

Uncertain

Posted
...

I am not sure including JS would have proven any more convincing. If he showed up with a high

probability I think the believing response would simply be the "real" authors are not included

in the analysis.

All the Best,

Uncertain

Probably so -- but there are plenty of other people around who are also interested

in the text, for various reasons (many of which relate directly to Mormon origins).

When we read Fawn Brodie and Dan Vogel, and are there emphatically told that

Joseph Smith, Jr. wrote the book, all by himself, with no other input than from the

KJV Bible and snippets from his private reading, then we just naturally look for

some quantitative analysis of the textual structure, to see if all of this can be true.

Various attempts at word-printing the BoM mainly arrive at the same conclusion --

that it is a composite text, written by more than one distinct authoring "voice,"

and perhaps subsequently edited, (which provides the text with a superficial

stylistic uniformity).

Whether we are trying to determine a Nephite authorship, authorship by Smith,

or some other sort of literary origin, I think it is important that the word-printing

continues to point us away from a single writer (and thus away from Smith as

the sole author).

Sans Brodie, the Tanners, Vogel and Metcalfe, we can all pretty much agree on

that conclusion -- so far as it goes.

Those who opt for a Nephite origin will be left to puzzle over authorship "voice"

distributions which do not agree with what the book itself says about its origins.

Those who opt for a 19th century origin will be left to puzzle over the possibility

of two or more writers cooperating (or unknowingly contributing) to a composite

text. Spalding-Rigdon advocates will at least begin to have an index as to what

and where in the text, each writer supposedly supplied his distinctive literary material.

What I hope for, in the future, are some additional studies that show distribution

of probable authors' input. To begin with, I do not much care whose names are

identified as the "authors;" be they Nephi, Mormon and Moroni, or Lucy Smith,

W. W. Phelps, and William Morgan,

Regardless of whose names we attach to those authorship distributions, we can

begin to learn something about the nature (and perhaps origins) of the text, by

comparing and contrasting the various patterns that show up, over a large

number of independent word-printing studies.

I hope that most of us can agree that at least that much analysis of the text is

useful -- no matter our quibbles over the meaning and relevance of methodologies

and seemingly subjective analytical outcomes/conclusions.

I, for one, want to have an answer in my shirt pocket, ready for questioners who

demand of me: "OK, Mr. Know-it-all; if Rigdon wrote part of the book, show me!"

Others may want an answer they can reliably hand out to skeptics who do not

believe Nephites could have written the book, etc. etc.

Can we all learn something useful from this process -- or is that expecting too much?

Uncle Dale

Posted
Probably so -- but there are plenty of other people around who are also interested

in the text, for various reasons (many of which relate directly to Mormon origins).

When we read Fawn Brodie and Dan Vogel, and are there emphatically told that

Joseph Smith, Jr. wrote the book, all by himself, with no other input than from the

KJV Bible and snippets from his private reading, then we just naturally look for

some quantitative analysis of the textual structure, to see if all of this can be true.

Various attempts at word-printing the BoM mainly arrive at the same conclusion --

that it is a composite text, written by more than one distinct authoring "voice,"

and perhaps subsequently edited, (which provides the text with a superficial

stylistic uniformity).

Whether we are trying to determine a Nephite authorship, authorship by Smith,

or some other sort of literary origin, I think it is important that the word-printing

continues to point us away from a single writer (and thus away from Smith as

the sole author).

Sans Brodie, the Tanners, Vogel and Metcalfe, we can all pretty much agree on

that conclusion -- so far as it goes.

Those who opt for a Nephite origin will be left to puzzle over authorship "voice"

distributions which do not agree with what the book itself says about its origins.

Those who opt for a 19th century origin will be left to puzzle over the possibility

of two or more writers cooperating (or unknowingly contributing) to a composite

text. Spalding-Rigdon advocates will at least begin to have an index as to what

and where in the text, each writer supposedly supplied his distinctive literary material.

What I hope for, in the future, are some additional studies that show distribution

of probable authors' input. To begin with, I do not much care whose names are

identified as the "authors;" be they Nephi, Mormon and Moroni, or Lucy Smith,

W. W. Phelps, and William Morgan,

Regardless of whose names we attach to those authorship distributions, we can

begin to learn something about the nature (and perhaps origins) of the text, by

comparing and contrasting the various patterns that show up, over a large

number of independent word-printing studies.

I hope that most of us can agree that at least that much analysis of the text is

useful -- no matter our quibbles over the meaning and relevance of methodologies

and seemingly subjective analytical outcomes/conclusions.

I, for one, want to have an answer in my shirt pocket, ready for questioners who

demand of me: "OK, Mr. Know-it-all; if Rigdon wrote part of the book, show me!"

Others may want an answer they can reliably hand out to skeptics who do not

believe Nephites could have written the book, etc. etc.

Can we all learn something useful from this process -- or is that expecting too much?

Uncle Dale

Hi Uncle Dale,

I agree with much of what say here. I do think it unfortunate if secular and religious scholarship proceeds on separate tracks rarely interacting in any substantial way. I personally think the word print studies provide ammunition to both sides. They have consistently shown multiple authors which is problematic for the JS did it crowd. And multiple authorship is not currently supported all that well by the historical evidence. Which leaves critics with a conundrum. On the other hand as you point out the distribution of voices does not fit all that well with what the text claims. In any case regardless of what this study does or does not prove. I think it is a step forward it would be nice to get more LDS scholarship in mainstream journals both pro, con and neutral.

All the Best,

Uncertain

Posted
...

it would be nice to get more LDS scholarship in mainstream journals both pro, con and neutral.

All the Best,

Uncertain

Agreed -- it is beginning to happen, but still has far to go.

UD

Posted

Hi Ben and Uncertain,

I agree with what looks like your concensus on the problems of calculating only a relative probability. What is left though, is a pattern of author identification that does not fit in with either an expectation of seeing Joseph Smith's imprint over the whole document, or of seeing a pattern matching claimed authorship for various Nephites. For example, large chunks of Alma are attributed to three different people (Rigdon, Cowdery, Spaulding) - testing only Chapters 1-44 when his son took over the plates.

Spaulding is the most interesting author, since there can be no claim he was influenced by the BoM later (as could be claimed with the Mormon authors). There should be no reason to see his word profile in the book more than Longfellow or Burrows.

In response to a question elsewhere, I had a tutu with MS excel to do a simple binomial probability test of 'Spaulding' against 'NotSpaulding'. Once you remove the 21 known Isaiah/Malachi chapters you are left with 218 BoM chapters, claimed by an assortment of Nephites. There is no reason any given Nephite should resemble Spaulding more than the others so a test of authorship would identify Spaulding about 1/7 times (about 14%) or 31 chapters (+/- 3 given the margin of error). [Actually I would expect the redacted Nephites to resemble redacted I/M more than a modern author, thus lowering Spaulding's chances, but I am being conservative here; also if the Mormons adopted a Nephite 'style' Spaulding's chances would be lower still].

Spaulding was identified as most likely author 52 times, or about 24% of the time which gave p=0.00005463. So, in spite of the problems of relative probability, the pattern of identification needs explaining. Why should Spaulding be identified disproportionately compared to Longfellow, Burrows, Pratt, and Cowdery?

Posted
Obviously neither LDS apologists nor critics like myself are going to budge on this... it will take more evidence. Maybe the word-print thing will move some people. DNA moved a few.

The Salamander Letter moved a few, too. Fortunately, the more that is learned the less effect it has. The DNA arguments have proven to be far less substantial than enemies of the church hoped and, unfortunately, it has failed to bolster LDS claims. As it stands, it's really a draw.

Some DNA theorists cried foul over the limited geography models that allowed for existing cultures to be in Mesoamerica at the time of Lehi's coming, but a close study of the Book of Mormon fully supports such models. And it's not like they were produced to answer the DNA charges; they've been kicking around for decades.

The thing I like about Spaldingâ??as opposed to a Smith only authorship theoryâ??is that it explains a lot of things the Smith only version doesn't (or at least the Spalding theory better explains)â??the large amounts of plagiarism being one.

Yes, but all the theories ignore the overall consistency of the Book of Mormon. As stated, the geography is astoundingly consistent, not only in Mesoamerica but in the Arabian deserts. We should not expect to see this sort of consistency in fiction. Cultural consistencies abound, too. Although we still haven't found a "Welcome to Zarahemla" sign in Reformed Egyptian (and we should be highly suspicious of anyone claiming to find anything in Reformed Egyptian), I would think that fitting the Book of Mormon into any sort of a geographic model in the Western Hemisphere would be like fitting a round peg into a square hole. One looks at the two "Bountiful" candidates (which are really comprised of a long strip separated by wilderness), and it becomes apparent that whoever wrote the Book of Mormon would have had to have highly detailed satellite images. With multiple writers, this problem of consistency would be greatly compounded.

"In this work he mentioned that the American continent was colonized by Lehi, the son of Japheth, who sailed from Chaldea soon after the great dispersion, and landed near the isthmus of Darien." - John Spalding, American Review, 1851

This would be a rather devastating blow if this could be proven. Change the quote from 1851 to 1829 and I'd be impressed. Show me the money.

Well I'm still learning all this, but John Spalding mentions: Now this, to me is interesting because we have Lehi colonizing Americaâ??but Lehi is the son of Japheth. Are we ever told in the BOM who is Lehi's father? As far as I know we are not. Next we have Lehi sailing from Chaldea, whereas the BOM also doesn't tell us that and LDS apologists apparently have settled on Yemen.

It most likely was in the 116-page manuscript (Lehi's father). Mormon apologists have settled on their "Bountiful" candidates because, if one actually goes to Arabia and uses only the Book of Mormon and a compass, that's where Lehi and his family would have ended up. It should have been a barren wilderness, but instead these areas are exactly as Nephi described them.

The point is...this is John Spalding, who we agreed is likely the most reliable of the Conneaut witnesses giving details we don't find anywhere else. How do we explain that?

Oh, come on. The testimony's date should give you the clue! The testimony later adds: "Lehi's descendants, who were styled Jaredites, spread gradually to the north, bearing with them the remains of antediluvian science, and building those cities the ruins of which we see in Central America, and the fortifications which are scattered along the Cordilleras."

At the time the Book of Mormon was published, those cities remained largely unknown. Much of the early speculation placed the location of the Book of Mormon events in New England, and Joseph Smith was largely unaware of those cities until drawings and descriptions of Mesoamerican cities such as Palenque were published and circulated by Frederick Catherwood in his 1842 work, Incidents of Travel in the Yucatan. That same year, Joseph Smith took an intense interest in the work and even went on record to say that Palenque was a Book of Mormon city. At first the scholars said that Palenque was too recent to be a Book of Mormon city, but then they discovered it was built on an even older city that did, indeed, date back to Book of Mormon times. John Spalding seems to be saying that his father wrote his romantic fiction as a way of explaining these ancient civilizations. The only problem with this was that they weren't generally known by Joseph Smith or his peers in the 1829-30 timeframe, nor did they seem to be common knowledge at that time.

Posted
...

Why should Spaulding be identified disproportionately compared to Longfellow, Burrows, Pratt, and Cowdery?

Good question -- there is something peculiar about both the amount

and the distribution of the "Spaldingish" word-print in the Book of Mormon.

Consider my composite chart of 1830 Alma XX-Helaman I, below:

plates6.jpg

(larger image here)

(entire BoM/Spalding chart, here)

The topmost graph indicates the level of words common to the Spalding's

"Roman Story" and each Alma page in the 1830 BoM.

The bottom bar-graph indicates the level of word-strings (phraseology)

common to the "Roman Story" and each Alma page in the 1830 BoM.

The middle, blue line is a generalized depiction of the Stanford researchers'

Spalding word-print levels (averaging the results of their two measurement

methods); it has mistakenly been labeled "non-contextual" on this graphic.

Why should four different types of quantitative testing yield the same

distribution pattern?

Possibilities:

1. Simple coincidence

2. An original Spalding authorship for this section of the BoM

3. Spalding was supernaturally influenced by the BoM text (magic)

4. Some other answer..... not yet discovered.

Your thoughts?

Uncle Dale

.

Posted

UD:

Thanks again for the reply.....

I'll present you with another possibility here ---

When the "Book of Lehi" was lost, due to Martin Harris' negligence, exactly what sort

of story about "Lehi" did those lost pages contain? Was it very much like the story told

in 1st Nephi? Or, was it something rather different?

That's the million dollar question! At the very least, according to Joseph Smith, apparently what we now have is, at least, much less specific and is, apparently, much more religious. It's also interesting how the book of Omni seems to be a sort of bridge between the two accounts, it's almost as though there was some need to get a bunch of names down on paper.

We have three ocean-crossing, American colonization stories told in the BoM:

1. The Jaredite account

2. The Lehite account

3. The Mulekite account

Now, which of those three stories does John Spalding's second statement best describe?

Which of those accounts is largely missing from the 1830 BoM text?

To be honest I get confused with all the talk of abridgments and small plates of Nephi, large plates of Nephi, plates of Mormon, Plates of Ether, plates of Brass, sealed plates, (Life on a plate), more particular accounts, etc. I think the loss of the 116 pages actually added to the complexity which ended up giving it a more realistically "abridged" feel. I think the Words of Mormon & Omni are a dead-giveaway to what's really going on, but excluding that, the loss was really a blessing in disguise because the end result looks more realistic due to the complexity. What I can't figure out is where Ether fits into the picture in relation to the lost portion.... was it an aborted first attempt that was later touched up and finished? Who knows?

What are the chances that John Spalding saw a different Lehite account than we now

have; and what are the chances that the Mulekite account was moved into 1st Nephi,

with a little "touching up" in order to make it fit well there?

Since you're suggesting it, I'd say probably pretty good!

If this conjecture is correct, the story told at the beginning of Spalding's manuscript

more resembled the Jaredite account than what we read in the Book of Mormon.

And the story of Zarahemla, as recalled by a few early witnesses, would then have

resembled the account we now have in 1st Nephi.

Is it possible to come up with an educated guess based on when each of the witnesses likely heard/read Spalding's manuscript? I'm guessing not due to too much missing info.

Just a conjecture --- but if you had to quickly replace 116 pages of storyline, and you

had the Mulekite account sitting in front of you, not yet "translated," what would be

your easiest way to come up with a replacement for the Book of Lehi?

Agreed, but let me play devil's advocate for a minute.... what about the fact that Smith was claiming he was actually giving a watered-down version of the same account? Or am I wrong about that? I always thought he was claiming to redo what was lost, only through a different (more religious) set of plates. No? What you're suggesting sounds to me like a totally different account with different people and events, correct? As far as I know, Smith was still thinking the missing pages might actually re-surface down the road sometime... so given that, wouldn't he have tried to keep the storyline as similar to the missing portion as possible rather than moving to an entirely different account?

Of course you'd have to alter passages here and there -- maybe even delete the

name of Jared's brother (Lehi?) to make things compatible throughout the book...

Yeah, that whole Brother of Jared thing is a mystery to me! I keep thinking there has to be a simple reason why he never gives the name.... ...wow, so you think the BOJ might have been Lehi? Why not just invent a new name, then? Lehoni... Lorantiam... the possibilities are endless and it's much easier than the mysterious "Brother of Jared." Who knows?!

I find that "loyal Mormons" put more stock in their own personal testimonies than facts anyway.

Sometimes that is the case -- in other instances they can be very much concentrated

upon knowing and making use of "facts." It all depends upon the person and situation.

Of course, you're right.

Agreed. I think the Rebecca Eichbaum testimony is compelling at putting Rigdon at Pittsburg at the right time (despite his denials) in light of the mail-waiting notice.

As I said before, even Rigdon's own relatives (one of them a Mormon)

tacitly admit that he was in Pittsburgh at an early date. That part of the

puzzle has long been solved.

Really? Who was that? LDS apologists don't accept it, do they?

My question, is where did Rigdon gain his

tanner's apprenticeship? And where did he sell leather goods during that

apprenticeship? In 1824-25 he was a journeyman tanner/leather dresser,

selling sheepskin bookbindings to Pittsburgh binders -- but he could not

have started out as a journeyman tanner. If he was selling leather to

Silas Engles in 1824, was he also selling leather to the Pattersons (who

were Engles' cousins in Pittsburgh and had a bindery) in 1816?

UD

I heard an MP3 file a while back that claimed he did sell leather bindings to Pattersons, but I don't know where they get the documentation to back that up... maybe it was just speculation.

Related question for you.... this may be way out on a limb but how likely do you think it is that Rigdon was Moroni?

Take care!

Posted

Steel:

The Salamander Letter moved a few, too. Fortunately, the more that is learned the less effect it has. The DNA arguments have proven to be far less substantial than enemies of the church hoped and, unfortunately, it has failed to bolster LDS claims. As it stands, it's really a draw.

Some DNA theorists cried foul over the limited geography models that allowed for existing cultures to be in Mesoamerica at the time of Lehi's coming, but a close study of the Book of Mormon fully supports such models. And it's not like they were produced to answer the DNA charges; they've been kicking around for decades.

Agreed, but I'd say the insertion of the word "among" in the BOM intro carries a lot more significance than it might appear to the casual observer. Limited geography theories aside, where would you place the Hill Cumorah?

The thing I like about Spaldingâ??as opposed to a Smith only authorship theoryâ??is that it explains a lot of things the Smith only version doesn't (or at least the Spalding theory better explains)â??the large amounts of plagiarism being one.

Yes, but all the theories ignore the overall consistency of the Book of Mormon. As stated, the geography is astoundingly consistent, not only in Mesoamerica but in the Arabian deserts.

I can grant that the most impressive evidence I've seen to support the BOM is in the Yemen connection, but when it comes to American geography, you simply have no BOM geography. There is no such thing.

We should not expect to see this sort of consistency in fiction. Cultural consistencies abound, too. Although we still haven't found a "Welcome to Zarahemla" sign in Reformed Egyptian (and we should be highly suspicious of anyone claiming to find anything in Reformed Egyptian), I would think that fitting the Book of Mormon into any sort of a geographic model in the Western Hemisphere would be like fitting a round peg into a square hole.

Why? If Nephites were real, then Nephite cities should be real. Cities don't just disappear. Reformed Egyptian should be a real language--even if its a dead one. Why is there no trace of it in Native languages? Why is there no known sample of it? If the BOM is true there must have been a heck of a lot of engraving going on over a really long period of time. Where did it all go?

One looks at the two "Bountiful" candidates (which are really comprised of a long strip separated by wilderness), and it becomes apparent that whoever wrote the Book of Mormon would have had to have highly detailed satellite images. With multiple writers, this problem of consistency would be greatly compounded.

That's a stretch. Maps were available in the 19th century. But I'm not even convinced the Yemen connection is valid to begin with.

"In this work he mentioned that the American continent was colonized by Lehi, the son of Japheth, who sailed from Chaldea soon after the great dispersion, and landed near the isthmus of Darien." - John Spalding, American Review, 1851

This would be a rather devastating blow if this could be proven. Change the quote from 1851 to 1829 and I'd be impressed. Show me the money.

If I could do that I could impress a lot of people! What you are admitting, though, is if John Spalding was telling the truth, there is a definite case to be made. I agree. The thing of it is, of course we're not going to see any pre-1830 quotes about a Spalding BOM connection because Spalding was not a household name until after the BOM was published and the similarities began to be noticed. The closest we have to that (that I am aware of) is testimony that Solomon Spalding suspected Sidney Rigdon of stealing his manuscript.

Well I'm still learning all this, but John Spalding mentions: Now this, to me is interesting because we have Lehi colonizing Americaâ??but Lehi is the son of Japheth. Are we ever told in the BOM who is Lehi's father? As far as I know we are not. Next we have Lehi sailing from Chaldea, whereas the BOM also doesn't tell us that and LDS apologists apparently have settled on Yemen.

It most likely was in the 116-page manuscript (Lehi's father).

Wait a minute.... if you are right then John Spalding was telling the truth. How could he have known about it otherwise?

Mormon apologists have settled on their "Bountiful" candidates because, if one actually goes to Arabia and uses only the Book of Mormon and a compass, that's where Lehi and his family would have ended up. It should have been a barren wilderness, but instead these areas are exactly as Nephi described them.

Like I said, the Yemen connection is the best evidence I've seen in support of the BOM, but I still think it's weak. Too much is made of the NHM connection. You want an 1829 date before you'll be impressed, and I want a group of secular archeologists to agree that Lehi traveled through Yemen.... then I'll be impressed.

The point is...this is John Spalding, who we agreed is likely the most reliable of the Conneaut witnesses giving details we don't find anywhere else. How do we explain that?

Oh, come on. The testimony's date should give you the clue!

The testimony later adds: "Lehi's descendants, who were styled Jaredites, spread gradually to the north, bearing with them the remains of antediluvian science, and building those cities the ruins of which we see in Central America, and the fortifications which are scattered along the Cordilleras."

At the time the Book of Mormon was published, those cities remained largely unknown. Much of the early speculation placed the location of the Book of Mormon events in New England, and Joseph Smith was largely unaware of those cities until drawings and descriptions of Mesoamerican cities such as Palenque were published and circulated by Frederick Catherwood in his 1842 work, Incidents of Travel in the Yucatan. That same year, Joseph Smith took an intense interest in the work and even went on record to say that Palenque was a Book of Mormon city. At first the scholars said that Palenque was too recent to be a Book of Mormon city, but then they discovered it was built on an even older city that did, indeed, date back to Book of Mormon times. John Spalding seems to be saying that his father wrote his romantic fiction as a way of explaining these ancient civilizations. The only problem with this was that they weren't generally known by Joseph Smith or his peers in the 1829-30 timeframe, nor did they seem to be common knowledge at that time.

No... you're not getting it... In fact you're lending support to my side! Think about what you're claiming.... you're saying that John Spalding obviously must have guessed right about the BOM "geography" in 1851 because by then Frederick Catherwood, etc, had published meso-American discoveries that even had Joseph Smith thinking the BOM geography might have included Central & South America. You think you're safe in that position, but you're not, because here's what John Spalding had said way back in 1833 while Joseph Smith was still giving revelations to the "Lamanites" in Missouri & Ohio:

They buried their dead in large heaps, which caused the mounds so common in this country. Their arts, sciences and civilization were brought into view, in order to account for all the curious antiquities, found in various parts of North and South America.

:P Oops.

The other thing is that John Spalding mentions things we don't find in the BOM. The claim--from your side--is that Hurlbut was implanting false memories in their brains from the BOM... John Spalding specifically mentions the "isthmus of Darien" whereas the BOM does not. He identifies Lehi's father as "Japheth" whereas the BOM does not. Those items were therefore NOT influenced by either Hurlbut or the BOM. There may be more where that came from, but I haven't studied it enough to find them yet.

Anyway, thanks for the reply, I appreciate your perspective on this.

Posted
I have some further comments for Uncertain. I spent some time reading up a bit on the subject here, and I have come to the conclusion that NSCC is probably very problematic in an application like this. I think that NSCC would work very well in a case where we have confident knowledge that the unknown author of a text was within a certain group of suspects. But without that confidence, it is quite useless. It might be useful even in taking apart a document with several known authors and identifying pieces that each was substantially responsible for. But, in this case, even with the anecdotal evidence provided in the essay, the identification of the group is rather weak - particularly when we note that they dismiss Joseph Smith entirely as a possible author.

The study actually seems to claim to be calculating the real probability that a particular author is the unknown author of a passage. However, this isn't what they are doing. There are several ways to show that this is the case. This might be the case if they were certain to have included all possible authors in the list of known authors (which is part of the reason for the otherwise non-essential historical narrative).

The first point I want to make is this - if we are using a system which can produce some kind of definite probability that a text with an unknown author A was really written by known author B, then we should be able to get the exact same results independant of what additional samples from other known authors are used - and more importantly under the condition where no samples from any authors other than author B are used. It should be apparent to anyone that this particular method, incorporating a single author with which to compare the unknown author too, will yield that author (Author cool.gif as the first choice for every sample of unknown Author A, every time. This is in part due to the nature of Nearest Shrunken Centroid Classification used (as discussed below).

The second issue deals with the method. The vocabulary used in the process (and thus the word-frequency matrix that is used as the basis for the Nearest Shrunken Centroid Classification - hereafter NSCC) is produced by involving all of the text samples from all known authors. This occurs twice. The first time it occurs is when a common vocabulary is built in which the vacabulary is common to every known author as well as to the unknown author. The second time it occurs is when this vocabulary is reduced by eliminating all common vocabularly elements that do not have a frequency across all of the samples of at least .1% - in other words, the additional authors tend to have a great deal of impact on the study. A clever response might be to talk about James Thurber, who wrote a book called "The Wonderful 'O'". This was a child's book written because of a bet with a friend, and in it, there are no words containing the letter 'o' (and thus the name). Imagine what including this as a sample text with a known author would do - the entire vocabulary would suddenly lose all words containing the letter 'O'. Now, obviously you wouldn't choose this text, but, as an extreme it illustrates the point beautifully. In effect, you can manipulate the vocabulary (and thus the matrix) by the selection of sample texts from known authors. This should also make it clear that the choices of texts and authors influences the study in some ways - meaning that it would be quite possible to get a different set of results simply by changing the make-up of the authors and texts - and in particular by changing the so-called control samples that get used. Since the process isn't time consuming at all, and since the authors of the essay admit that their choice of control texts was deliberate and not random in any way (based on certain features which, as I noted, otherwise seem to have no impact on the study at all), I actually do suspect that these were chosen for the purpose of influencing the outcome of the study.

A third aspect of this comes now from the nature of NSCC that they use. If we go to the website which they link to in the footnotes where the complete description of the process can be found (and there is a lot of interesting things to read there), we very quickly get to things like this:

http://www-stat.stanford.edu/~tibs/PAM/Rdist/howwork.html

"Nearest centroid classification takes the gene expression profile of a new sample, and compares it to each of these class centroids. The class whose centroid that it is closest to, in squared distance, is the predicted class for that new sample."

Do you see how this works? It compares centroids (which is a way of condensing the matrix of frequencies of disparate elements) of a known set and tells you which of the known set an unknown sample is "closest to". This is not an absolute. This is a relative comparison. It means that the results are always determined by the starting positions of the samples. When doing cancer research, this is a great tool because we have a relatively finite number of cancer types we would be using this to genetically type, and our concern is to discover which one it is closest to as quickly as possible using a wide range of data (the mapped genes) and comparing huge amounts of data very quickly. It also couldn't work with a single cancer type comparison - because, if we only have X as a sample, and we ask which sample unknown type Y is closest to, of course all we get is X. Within the cancer research, where there is a reasonable assumption that the cancer being studied is indeed one of the several types it is being compared to, this can be tantamount to a non-relative probability (that is, if we are certain that the actual identity of the type is in the set, and the set contains all possible types, then this could be a rather definite probability rather than a relative one).

In terms of demonstrating this flaw, all that has to be done (and it will take some time for me - at least a week or so), is to add some samples to the composite list - or to generate a new set of samples - and show that changing the sample sets produces a wildly variant conclusion or showing that a completely different set yields similar results with different candidates - and thus showing that there is no value to the conclusions being drawn. Since the essay seems to assume from the beginning a narrow list of possible authors who must (exclusively) be the authors of the text, this may be a logical way to deal with that question, but to me, it doesn't seem reasonable at all. And if we were to add another author to the list of authors sampled (say from my list, or by taking some of what we have from Joseph Smith), we might well get wildly divergent results. And certainly, were we to change the control texts we might find something else as well.

There are several things I would have liked to have seen. I would have liked to have seen a set of control text (several known authors) showing what they suggest we should see from a series of unconnected authors and texts - a random pattern with roughly the same number of samples being assigned to each author. I think we would see patterns of authorship that are clearly fabrications (and nothing like their figuring random distribution which it seems that they do).

Ben M.

Ben, you have some excellent criticisms here, but it seems to me you are a little too eager to throw the baby out with the bathwater. Clearly there are a number of weaknesses that increase the uncertainty of interpreting the results and certainly doesn't prove a Rigdon-Spalding connection to BoM origins, but likewise there is still a great deal of information about possible authorship that can be gleaned from this study. I posted a comment similar to the following on MDB pointing to where I think we can really see what information about authorship is being extracted by the applied centroid method:

First of all, what should we expect the results to be assuming the BoM is of Nephite origin? The answer depends strongly on whether the translation was 'loose' or 'tight'. A loose translation should have definite imprints of a strong Joseph Smith signal across the entire work. A tight translation should result in no overwhelming JS signal or any overwhelming signal correlating with any of the selected author pool, giving a flat probability distribution function across the board (not counting Isaiah-Malachi which will clearly show up in either as a special case).

Now, Joseph Smith wasn't included, so that should leave us to consider what a JS signal should look like. My supposition is that a JS signal for this group of authors would either be flat (i.e., little correlation with any and mapping everywhere simultaneously) or would preferentially map on Cowdery as I would guess he would be more similar to Smith than any of the others. Another necessary and important note to mention is that errors in data selection and such should again bias the results to a flat distribution (more controls relative to actual signal).

So what does it mean that the results are not flat? For starters, that seems to me to rule out a tight translation (I doubt many here would concern themselves with that). So what to with the strong Rigdon-Spalding pairing? Maybe one could suppose that the unaccounted for JS signal preferentially mapped there, indicating a loose translation. Personally, that particular pairing is a little too striking for me to chalk up that way, and seems that I should take the Rigdon-Spalding hypothesis more seriously than I have in the past. Clearly, not the only way they must be interpreted (especially in the absence of a Smith centroid), and it will be interesting to see how robust the signal remains in future studies, particularly when Joseph Smith can be reliably included. But at the moment, we have a quantitative result that points to the Rigdon-Spalding connection of BoM origins that has a great deal more weight to it than most here seem willing to grant.

Posted
In response to a question elsewhere, I had a tutu with MS excel to do a simple binomial probability test of 'Spaulding' against 'NotSpaulding'. Once you remove the 21 known Isaiah/Malachi chapters you are left with 218 BoM chapters, claimed by an assortment of Nephites. There is no reason any given Nephite should resemble Spaulding more than the others so a test of authorship would identify Spaulding about 1/7 times (about 14%) or 31 chapters (+/- 3 given the margin of error). [Actually I would expect the redacted Nephites to resemble redacted I/M more than a modern author, thus lowering Spaulding's chances, but I am being conservative here; also if the Mormons adopted a Nephite 'style' Spaulding's chances would be lower still].

Spaulding was identified as most likely author 52 times, or about 24% of the time which gave p=0.00005463. So, in spite of the problems of relative probability, the pattern of identification needs explaining. Why should Spaulding be identified disproportionately compared to Longfellow, Burrows, Pratt, and Cowdery?

Hi Danna,

I ran the numbers in your post and got a p-value of 5.7e-5 so pretty much what you got. This is not a surprise Jockers et al. using a more sophisticated statistical hypothesis test demonstrated the pattern of authorship attribution is highly non-random. So showing that Spaulding shows up as the author (out of 7 authors) of BOM chapters more than would be expected by random chance is not surprising. I do think the non-uniform distribution of authorship is not what you would expect if the set of authors had nothing to do with the BOM. The key question is why this is the case. Taking a quick glance at Uncle Dales graph he posted in post #166 it looks to me like Spaulding shows up as the author most resembling the given passage when the passages are more narrative and not theological in content. This could be because the set of texts used to construct Spauldings word print are more narrative in origin. Hence it would not be surprising those passages show up as closer to Spauldings word print. If I was an apologist that is the first thing I would look at.

All the Best,

Uncertain

Posted
...

Why should four different types of quantitative testing yield the same

distribution pattern?

Possibilities:

1. Simple coincidence

2. An original Spalding authorship for this section of the BoM

3. Spalding was supernaturally influenced by the BoM text (magic)

4. Some other answer..... not yet discovered.

Your thoughts?

Just to clarify the generalized blue graph line in my posted BoM text graphic,

it was derived from the Spalding Word-Print Chart, Part A, in the Stanford paper.

Here's my excerpt (40% and above) from that paper's Spalding chart, Part A:

SpaldMap8.gif

(larger image)

(My 1980 chart of entire BoM)

Plates2.gif

Compare the concentration Spalding voice/signal occurrences in the latter part

of Alma, with my own, less scientific mapping out of "Spaldingish" (red) text there:

plates4.jpg

Query: Would a computer analysis overseen by LDS statisticians yield significantly

different results? Would a "proper"analysis and mapping, ala Ben's suggestions,

probably eliminate this remarkable Spaldingish cluster in Alma?

Uncle Dale

.

Posted
... according to Joseph Smith, apparently what we now have [in 1st Nepji] is, at least, much less specific and is,

apparently, much more religious. It's also interesting how the book of Omni seems to be a sort of bridge between

the two accounts, it's almost as though there was some need to get a bunch of names down on paper.

My guess is that 1st Nephi cannot read too differently than what

the Book of Leji read like -- or the original scribe, Martin Harris,

would have probably offered up some vocal complaints.

On the other hand, 1st Nephi could be lacking a great deal of

story-line that was present in the Book of Lehi. Deletions and

additions would not have aroused Martin's objections, if the

replacement story was generally compatible with the earlier text.

let me play devil's advocate for a minute.... what about the fact that Smith was claiming he was actually giving

a watered-down version of the same account? Or am I wrong about that? I always thought he was claiming to

redo what was lost, only through a different (more religious) set of plates. No? What you're suggesting sounds

to me like a totally different account with different people and events, correct?

What I am saying is that a story of ten tribes wandering across Asia and

coming into Alaska, or some other part of North America might have been

eliminated, and with that elimination, John Spalding's recollections would

no longer match the replacement story.

The Mulekite colonization account is largely lacking -- the story of a group

of Israelites (mostly proto-Jews) who set out from Jerusalem shortly before

its fall to the Neo-Babylonians, and crossed the sea to the Americas, to

establish Zarahemla. THAT story is not much different from 1st Nephi.

As far as I know, Smith was still thinking the missing pages might actually re-surface down the road sometime...

so given that, wouldn't he have tried to keep the storyline as similar to the missing portion as possible rather

than moving to an entirely different account?

Yes -- there would be that constraint. So a totally new Lehite account

could not have been fabricated -- the basic characters must have been

the same. But much could have been eliminated from the first story,

and additional details could have been moved up from the now missing

Mulekite story.

Yeah, that whole Brother of Jared thing is a mystery to me! I keep thinking there has to be a simple reason why he never gives the name....

...wow, so you think the BOJ might have been Lehi? Why not just invent a new name, then? Lehoni... Lorantiam... the possibilities are

endless and it's much easier than the mysterious "Brother of Jared." Who knows?!

Brother of Jared is perhaps copied from "Brother of Shared" in the same book. But

the elimination of his name coincides with what John Spalding recalled of a Lehi

journeying to the Americas in what is otherwise the Jaredite story. Coincidence?

"As I said before, even Rigdon's own relatives (one of them a Mormon)

tacitly admit that he was in Pittsburgh at an early date. That part of the

puzzle has long been solved."

Really? Who was that? LDS apologists don't accept it, do they?

It's in John E. Page's re-write of Winchester's anti-Spalding pamphlet, at

my Spalding web-site. I think LDS apologists have to accept it -- at least

the RLDS do (they even re-printed Page's pamphlet and sold it).

http://www.solomonspalding.com/docs/Page1843.htm#1866

Related question for you.... this may be way out on a limb but how likely do you think it is that Rigdon was Moroni?

Whitsitt thought that they were one in the same -- also the same as John the Baptist.

Read Pratt's "Angel of the Praries" for something similar.

As for me, I'm reluctant to read too much into the Moroni story. He was originally

the Angel Nephi -- "angels" in Alexander Campbell's translation of the Bible, are

merely "messengers" and "fellow servants" in spreading the gospel -- exalted men,

perhaps. Or, Taking Campbell/Rigdon very literally, any messenger is an "angel."

UD

Posted
Why not? Why would you ruin your own credibility and the credibility of the Book of Mormon as scripture by outing yourself as its forger? It seems to me that if Rigdon and Spalding believed strongly enough in the message of the Book of Mormon to go to such trouble to bring it into the world, they wouldn't be particularly inclined to expose it regardless of its partial co-option by Joseph.

We also have other witnesses to the book of mormon who seemed to have not exposed the 'fraud' when they all had the chance. This makes it all the more problematic. And also a reason for the witnesses, so that such discussions like this would not be a slam dunk for the fraud explanation.

Posted
Good question -- there is something peculiar about both the amount

and the distribution of the "Spaldingish" word-print in the Book of Mormon.

Consider my composite chart of 1830 Alma XX-Helaman I, below:

plates6.jpg

(larger image here)

(entire BoM/Spalding chart, here)

The topmost graph indicates the level of words common to the Spalding's

"Roman Story" and each Alma page in the 1830 BoM.

The bottom bar-graph indicates the level of word-strings (phraseology)

common to the "Roman Story" and each Alma page in the 1830 BoM.

The middle, blue line is a generalized depiction of the Stanford researchers'

Spalding word-print levels (averaging the results of their two measurement

methods); it has mistakenly been labeled "non-contextual" on this graphic.

Why should four different types of quantitative testing yield the same

distribution pattern?

Possibilities:

1. Simple coincidence

2. An original Spalding authorship for this section of the BoM

3. Spalding was supernaturally influenced by the BoM text (magic)

4. Some other answer..... not yet discovered.

Your thoughts?

Uncle Dale

.

I have been down this road before with these graphs. If I remember correctly Warship made a comparison with some sort of book but I forgot the name of that book which also showed a correlation with book of mormon wording. It tended to negate your graphs. Was it the popa val or something sounding like it?

Posted
I think that investigators like myself will have a better chance of

talking to the interested non-Mormons, now that the Stanford team's

paper has been published.

We will never convince Sandra Tanner or Dan Vogel of its relevance

in investigating Mormon origins -- but there is a new generation of

non-Mormons (and former Mormons) on the rise, who may be more

susceptible to Spalding-Rigdon authorship explanations.

Trying to convince the LDS and RLDS of such things is a lost cause;

but the younger Gentiles may now take the trouble to listen -- and

to edit future Wikipedia articles.

UD

It is not a lost cause if hypothetical illustrations are replaced with convincing evidence. However, no logical explanation exists as to why so many people would be in on the fraud, have a fallen out with Smith and still stick to the story line that the book came from gold plates. And this would include Rigdon, and Cowdery and David Whitmer, not to mention Emma during her hatred for everything containing polygamy. It seems to fly against human nature not to bring a 'fraudster' down in degrace when the chance showed itself for the witnesses and Rigdon.

Archived

This topic is now archived and is closed to further replies.

  • Recently Browsing   0 members

    • No registered users viewing this page.
×
×
  • Create New...