Jump to content
Seriously No Politics ×

Two voices in the Book of Mormon


robuchan

Recommended Posts

Posted

Yep

As a layman on this subject, I confess I stared at that thing for about 10

minutes trying to figure it out before I gave up.

I assume that in its final, published form, it is accompanied by a data

tabulation indexed to all the textual segments plotted on the graph --

(such as 16 Spalding segments, 235 BoM segments, etc.)

But, without some descriptors for the four quads of the pca plottings,

I'm not quite sure what deviations from the mean are being measured.

Perhaps Bruce will open up any pass-word guarded pre-pub stuff, and

let us have a gander at the text of the paper... (fingers crossed).

UD

Posted

A quick question:

What edition of the Book of Mormon have been used for the wordprint studies? It seems to me that if they used the earliest text reconstructed by Royal Skousen they would get different results.

Posted

A quick question:

What edition of the Book of Mormon have been used for the wordprint studies? It seems to me that if they used the earliest text reconstructed by Royal Skousen they would get different results.

I'm not sure which version Bruce has made use of.

Jockers and associates used the published 1830 version,

broken into segments, based upon the modern LDS chapter

divisions.

When it comes to measuring the occurrence of frequently used

non-contextual words, for word-print studies, I doubt that

the Skousen reconstruction differs by more than half a percent

from the 1830 edition.

"Close enough for government work."

UD

Posted

Are you familiar enough with the representation of data in principal component analysis charts, to explain to us in layman's language exactly what we are seeing, when the majority of BoM chapter plots cluster in the x-positive and y-positive quad of deviations from the mean?

I suppose I can try. But I am not that good at explaining statistical modeling.

For a discussion of the math behind this representation, you can go here: http://en.wikipedia.org/wiki/Principal_component_analysis

Bruce quoted from this - but even there its a bit hard to understand.

For the moment, lets look at the chart you offered - it was a rather flat graph. Easy numeric statistics making up a nice X-Y coordinate access.

In the case of this particular study, we have something like 130 different vectors for each piece of text. (The vectors identify the vocabulary word and a frequency associated with it). Now, we can't really make a picture showing all 130 different vectors for each of the chapters in the text. Not only would it be difficult to look at, if we can't evaluate it in some fashion, its nearly impossible for us to actually use by itself. So, PCA gets used to create a modeling of the data.

So we get these principal components. By principal component, what happens is that the data is taken and the projection of the data with the greatest variability becomes the first principal component. Then, the next principal component is the second most variable projection and so on. You can actually map as many of these as you want - but of course, visually, you run into problems once you get beyond 3 variables (we can't really visualize a fourth dimensional space very well). However, since you have covered the most variable projections, you have created a low resolution image of your composite data - what Bruce called a "shadow" of the data. You can get a decent feel for the distribution of all of your 130 data points in this way - and do so in a fashion that covers the largest variability. So, to use a chart like this, you express the most information that you can about your data in a simple X-Y coordinate chart. Some information is lost - and going to a three dimensional chart might recover some of what is lost - but, even with some lost information, the process is remarkably accurate and useful for what it does. You will notice that the greatest variation occurs along the X axis. The Y axis has some variation, just not as much. Were we to add a Z axis, it would provide less variation than the Y axis. So, each dimension we might add on adds some data, but that data is progressively smaller. We might see a standard deviation or two on the Z axis - but no more than that. And, in a sense, this would simply increase the distance to the mean, not decrease it.

Ok, so back to the chart. In PCA descriptions like this, the data is mean centered. This is useful for highlighting comparisons in situations like this. We aren't so concerned with which direction is which - in this case, we are interested in the distance from the data point to the center (greater distance means greater variability from the mean). Since we are working with two principal components, points that are in opposite corners are much farther apart than objects that share one side or the other of the chart.

So, looking at this chart - what we notice first of all is that the bulk of the chapters of the Book of Mormon - particularly those in the positive X and positive Y quadrant - are no where near the other test authors. That is, almost everything is on the other side of the center (the mean) from these data points.

Now, on the left side, where it starts to merge (the data points begin to cluster up near the Book of Mormon), the Green 'X's are the data points for the Isaiah chapters. So, where we see Book of Mormon data points in there, we expect that these are the Isaiah chapters in the Book of Mormon. And you will notice that the samples from Rigdon that the Jockers study used are quite close (relatively speaking) to both the Isaiah author and the Book of Mormon Isaiah chapters. The Spaulding author, for example, is quite close to Pratt - and is clearly (from the chart) closer to the Book of Mormon in places than the other authors - the Spaulding author isn't actually very close to the Book of Mormon text at all - certainly not close enough to assume some kind of correlation. The Barlow author is about as far away as you can get - except for some outlier points (which are still some distance away). You can see from this chart why, perhaps, the Barlow control author doesn't work very well. If we are expecting Barlow to function as some kind of control - we have a problem. The Barlow text is so distant that it cannot really help us in this case. And you can see from this plot that there are a couple (less that a half dozen) BoM chapters which might be associated with Rigdon, which are close enough to be considered. But at the extremes, even when two points are 10+ standard deviations apart, the Stanford study still identifies a potential author - and may assign a fairly high confidence level.

So, realizing that this might still be greek, does it help at all?

As far as the Book of Mormon text, I believe Bruce used the same text that Jockers used.

Ben.

Posted

...

So, realizing that this might still be greek, does it help at all?

...

Yes, that was an excellent explanation. I now understand why the

four quads were not labeled -- because there are more than four

variables involved.

Whatever the factors creating clusters in any one part of any particular

quadrant may be, they are not communicatable in simple terms.

A few BoM chapters plot within the cluster of Spalding chapter plots,

and several more BoM chapters plot at distances away from a Spalding

chapter less than the greatest distance between two Spalding plots.

I take that fact to represent that, in some way, those BoM texts

resemble Spalding texts. As you might guess, I'm curious to know

which parts of the BoM those might be.

One last question --- If only Spalding and the BoM texts were plotted

on such a cluster graph, do you suppose that the resultant visual

representation would closely resemble what we see in the multi-text

graph? Or, would the elimination of Rigdon, Isaiah, Longfellow, etc.,

cause such a great change in the equations/relationships, that radically

different plottings would result?

UD

Posted

Ben, do I understand you to say that this chart shows only the two word vectors with the largest variability across all sample texts? (As opposed to Delta, which averages the absolute values of the differences between texts for all vectors?)

Posted

Yes, that was an excellent explanation. I now understand why the

four quads were not labeled -- because there are more than four

variables involved.

Sorry to drop off the radar. We had a birthday and a baptism for a grandchild - a great day! Didn't even think about the Rigdon-Spalding theory.

One idea I might add to Ben's excellent description of PCA is to think of an overhead projector. An overhead projector produces a 2-dimensional shadow of a 3-dimensional object held over the projector. If you held, say, a pencil over the projector you could see anything from the image of a dot to the image of an arrow, depending on how you rotated the pencil. The image of the longest possible arrow would be the PCA plot of the pencil.

The authorship data is a cloud of 456 points (one for each text) that lives in 110 dimensional space (one dimension for each word). The 2-dimensional PCA plot is like using a mega-super-duper overhead projector to view this point cloud in 2 dimensions. The 2-dimensional PCA plot is the shadow of the 110-dimensional cloud that shows the greatest spread.

Note that the authorship information is not used in producing the PCA plot. If it so happens that the PCA plot is such that the texts are clustered by author, we've learned something pretty important about the structure of the data.

Does this help, or does it confuse things more?

A few BoM chapters plot within the cluster of Spalding chapter plots,

and several more BoM chapters plot at distances away from a Spalding

chapter less than the greatest distance between two Spalding plots.

I take that fact to represent that, in some way, those BoM texts

resemble Spalding texts. As you might guess, I'm curious to know

which parts of the BoM those might be.

The PCA plot in my oringinal comment uses Criddle

Posted

Ben, do I understand you to say that this chart shows only the two word vectors with the largest variability across all sample texts? (As opposed to Delta, which averages the absolute values of the differences between texts for all vectors?)

The axes of the 2-dimensional PCA plot are the 2 linear combinations (or composites) of all features (relative frequencies of 110 specific words in this case) such that maximum variation among the texts is displayed graphically.

PCA is just a convenient way to visualize high-dimensional data. It's not an alternative to Delta, which is a methodology for attributing texts of unknown authorship to one of a specified closed set of candidate authors. However, it is a useful tool for determining if Delta is appropriate, or if it produces nonsense. If PCA shows that the texts of unknown authorship are completely distinct from any of the candidate authors, Delta (and NSC) is/are useless.

One other thing. If the texts cluster by author in the 2-dimensional PCA plot, the distinction among the clusters can only increase as you consider more principal components. So if you see a difference (say between the Spalding texts and the BOM chapters) in the 2-dimensional PCA plot, that difference is not misleading. The difference can only become more pronounced as you consider more components.

Posted

Noel00,

I see that you are reading this discussion board.

I would like to ask Craig why he has never (to my knowledge) attempted to answer the main criticism of his work, namely that his authorship attributions are meaningless since his closed-set NSC method is forced to pick a winner out of the candidates he included. Jeff Lindsay made this argument eloquently on his blog (Mormanity), Ben Maguire put forth the same argument on this thread , and other bloggers (whose names I can

Posted

I sent the latest comments to Craig. You have to cut him some slack , he has been engaged in battle in three fronts, Chris Smith, Don Bradley and Maguire and the latest will be Bruce. He once told me that the church with unlimited resources and staff at BYU would soon join the fray. He said he has some deadlines to meet at the moment but will be back.

Posted

...

He said he has some deadlines to meet at the moment but will be back.

I very strongly suspect that Craig's two associates in the authorship

attribution project are advising him that the best place to reply to

critical responses to their research would be in the professional

literature.

If Bruce's work is as relevant and substantial as it appears to be,

then I also suppose that his reporting will eventually appear in

the literary/linguistic/computing peer-reviewed literature -- and

perhaps also in the statistical/mathematical literature. My guess

is that any meaningful reply from the Stanford team will also appear

in such journals.

But, short of a formal citation and counter-criticism, defense,

further exploration, etc., I would be very curious to hear something

of the Stanford team's characterization of the pca plotting of

their own data.

Since Bruce's chart was derived from that team's own data, the team

members should be able to reverse-engineer a mathematical process

which identically duplicates Bruce's results.

In such a re-run of the process, the Stanford team should be able to

succinctly identify the text segments plotted on the chart, and thus

isolate clusters of attributed/known authorship within the depiction.

For example, identification of the "Moroni" plots may (or may not)

allow for the highlighting of a close cluster of Moroni chapters,

indicative of a common authorship (and perhaps a single authorship).

On the other hand, were the various plots of Moroni chapters to fall

into widely separated regions on the chart, that phenomenon might

also tell us something important.

Hopefully one of the Stanford team will endeavor to duplicate and

diagnose Bruce's charted findings, and will then have something to

say to an interested audience (us). I'm happy to wait a couple of

years to see their formal response appear in the literature.

I've grown used to waiting in this line of investigation.

UD

.

Posted
He once told me that the church with unlimited resources and staff at BYU would soon join the fray.

For purposes of clarification:

"The Church" will not "join the fray."

I've never received a directive from "the Church," nor from anybody in the leadership of the Church, to enter into any particular issue. Nor, so far as I'm aware, have any of my associates.

It would be nice if we had "unlimited resources and staff" behind us, but we don't.

This particular matter has been no different from any other, and there is no reason to believe that things will change.

Craig Criddle should lay not that flattering unction to his soul.

Posted

...

Note that the authorship information is not used in producing the PCA plot. If it so happens that the PCA plot is such that the texts are clustered by author, we've learned something pretty important about the structure of the data.

Does this help, or does it confuse things more?

...

No confusion at all now, Bruce. You've articulated all the main points, I think.

I assume that if the process were reversed (charting closest non-variations from

the mean), that we'd simply end up with a huge cluster around the 0,0 point?

If you've followed the ongoing discussion here, over a considerable period

of time, in many threads, you'll recall that the vocal advocates for sundry

BoM origin explanations fall into three general categories:

1. Jaredites and Nephites wrote it; Smith translated into his vernacular

2. Smith wrote it by himself (incorporating a little KJV material)

3. Smith also incorporated some non-biblical pre-existing texts

I'm not sure just how much your chart can help us here -- to better

define and "rate" those three possibilities -- but you can imagine my

interest in the question.

Do you have any suggestions as to how your findings might be incorporated

into future investigations, of whether the BoM text is attributable to

a single author (probably Smith) or to more than one author (either ancient

or modern)?

If so -- you may soon find yourself swamped by questions. This can be

interesting stuff, when the focus of the discussion begins to fall

within the scope of our respective personal belief systems.

Uncle Dale

Posted

I sent the latest comments to Craig. You have to cut him some slack , he has been engaged in battle in three fronts, Chris Smith, Don Bradley and Maguire and the latest will be Bruce. He once told me that the church with unlimited resources and staff at BYU would soon join the fray. He said he has some deadlines to meet at the moment but will be back.

Noel00,

I have never met Chris Smith, Don Bradley, or Ben Maguire (although I think I

Posted

Since Bruce's chart was derived from that team's own data, the team

members should be able to reverse-engineer a mathematical process

which identically duplicates Bruce's results.

In such a re-run of the process, the Stanford team should be able to

succinctly identify the text segments plotted on the chart, and thus

isolate clusters of attributed/known authorship within the depiction.

.

I know they could. In fact, they did provide PCA plots for their recent work with the Federalist papers (see here). If you look there, you will see that the Federalist Papers application is an ideal place to properly apply their NSC methodology. The points representing the unknown texts are completely among texts of the candidate authors in the PCA plots. I just wonder why they didn't supply such plots for thier work with BOM chapters.

Posted
If you look there, you will see that the Federalist Papers application is an ideal place to properly apply their NSC methodology.

That was my conclusion, as well. Federalist Papers, yes. Book of Mormon, no.

Posted

That was my conclusion, as well. Federalist Papers, yes. Book of Mormon, no.

In his recent paper Jockers says: "The Federalist Papers corpus was

selected for this research on the grounds that it met the two primary criteria of

being both familiar to authorship researchers and of adequate size to afford

thorough testing. As pointed out by a reviewer, the Federalist corpus is not the

only suitable problem set for a benchmarking analysis. However, in addition to

being one of the most widely used corpora for authorship attribution testing and

investigation, the Federalist Papers is a "real" authorship corpus, with an ample

set of works of known authorship and a smaller subset of disputed texts. It has

the advantage of being well understood. As early as 1997, Richard S. Forsyth

had noted that the Federalist Paper problem 'is possibly the best candidate for

an accepted benchmark in stylometry'..."

Do you suppose that the Stanford team members have violated their own

conclusions here, by admitting the Book of Mormon into the corpus of

literature examinable by methods applicable to this "benchmark in stylometry?"

If so, then what automated methodology is applicable to the study of differences

in language use, throughout the various literary sections of the BoM text?

UD

.

Posted

Wouldn't it be your previous suggestion of comparing the three most Spalding-like units with the rest of the text? The more I think about it, the more I like it.

And yes, In case you didn't know it, I am VERY impatient.

Just a drive-by. :P

Posted

Wouldn't it be your previous suggestion of comparing the three most Spalding-like units with the rest of the text? The more I think about it, the more I like it.

And yes, In case you didn't know it, I am VERY impatient.

Just a drive-by. :P

If this pca charting (or other useful methodologies) can isolate some

Alma chapters "relatively close" to a cluster of known Spalding texts,

then THAT discovery would provide me with one more criterion by which

to select the three or four "most Spalding-like units" from the BoM,

for cross-comparison with the entire BoM text.

It is an experiment that I'd like to see conducted, but so long as we

are having this controversy over textual categorization methodologies,

I suppose we ought to hold off, and wait to see what else develops, in

helping identify those "most Spalding-like units."

For example -- let's say we identified all the BoM chapter plottings on

Bruce's chart, and discovered that some of the chapters from the latter

portion of Alma were among those "relatively close" to the Spalding texts.

By "relatively close," I mean not more than two or three standard deviations

distance from the center of the "Spalding cluster."

At first glance, such distant BoM chapters might be viewed as having

nothing in common with Spalding's use of language in his own writings.

But, in comparison to the remainder of the plotted BoM chapters, those

"relatively close" points would probably resemble Spalding, more so

than any other part of the Nephite Record.

What would catch my attention, is if such "relatively close" BoM plots

represents chapters from the latter part of Alma, which we have already

identified as resembling Spalding's vocabulary and phraseology.

If we can get to THAT point in our investigations, then I'd be very

curious to compare those "most Spalding-like units" to the rest of the

BoM, to see what new (hitherto unrecognized) correspondences might show up.

But there is an entirely different use I'd like to set up for the pca plots,

and that is an internal examination of the Book of Mormon chapters themselves.

Let's say that all of the chapters comprising Moroni plotted out in a

tight cluster, along with the Moroni portions of Ether -- but that the

remainder of the Ether chapters plotted out far away on the pca chart.

Such a discovery would support the idea that Moroni wrote the book

attributed to him, and added in the editorializing in Ether (whoever

"Moroni" might have been).

But -- to go on a little more -- what if we saw a run of 2nd Nephi

chapter plots ALSO cluster with the Moroni chapters? That would be a

very strange discovery, since Moroni the son of Mormon should not have

been writing anything on Nephi's plates.

There are a lot of experiments we might conduct here. But first of all,

we need to identify all of those unlabeled BoM plots on Bruce's chart.

UD

.

Posted

Yeah, I have a lot of questions about Moroni (except for ch. 9)

It would be worthwhile to play around with that, and see what else it resembles.

I'd better dodge out of here, before I get stalked by the dinosaur.

Posted
Do you suppose that the Stanford team members have violated their own conclusions here, by admitting the Book of Mormon into the corpus of literature examinable by methods applicable to this "benchmark in stylometry?"

The Federalist Papers are appropriate for this methodology, but not for the reasons cited by the Stanford team. To quote again from my paper,

"Although the Federalist Papers are a classic case, however, they are not really comparable to the Book of Mormon. According to Shlomo Argamon, the assumptions of word-frequency analysis 'fundamentally limit use of the method [to cases in which] all the samples (from all authors) are of pretty much the same textual variety, otherwise we would expect the word frequency distributions over the comparison set to be a mixture of several disparate distributions, one for each genre found in the set, thus potentially biasing results depending on the variety of the test text.' The Jockers and Witten study of the Federalist Papers satisfied this criterion, but the Jockers, et. al. study of the Book of Mormon unequivocally did not."

And in a footnote I add,

"Even Jockers and Witten admit that 'context-specific words' can adversely impact the results. Thus even in the Federalist Papers case, 'if the Madison training texts and the test texts address a particular topic that is not addressed by the Hamilton or Jay training texts, then the NSC classifier might use these words as very strong evidence that the test texts were written by Madison.'"

Another footnote further explains,

"The application of Delta to the Book of Mormon is further complicated by studies which show that when authors who have no familiarity with statistical attribution methods attempt to obfuscate their style or to imitate the style of another author, the accuracy of stylometric methods is reduced 'to the level of random guessing.' Since the Book of Mormon imitates the King James Version of the Bible, stylometry is unlikely to be useful in determining its authorship. See Michael Brennan and Rachel Greenstadt, 'Practical Attacks Against Authorship Recognition Techniques,' available from www.cs.drexel.edu/~greenie/brennan_paper.pdf [accessed April 16, 2010]."

If so, then what automated methodology is applicable to the study of differences in language use, throughout the various literary sections of the BoM text?

If I knew the answer to that, I'd have used it already. I'm not sure there is a perfect methodology for this purpose, but then I'm not an expert in statistical authorship attribution.

Peace,

-Chris

Posted

...

Since the Book of Mormon imitates the King James Version of the Bible, stylometry

is unlikely to be useful in determining its authorship.

...

Perhaps so -- and yet, by automated means is appears possible

to identify the Isaiah chapters, from out of the text as a whole.

Unless I'm totally misreading what Bruce has communicated, it seems

that something so common as pca charting can separate out the Isaiah

chapters, and cluster them away from the remainder of BoM plots.

If that is possible, then I'm guessing that the "Moroni" voice could

be similarly separated from most of the plots, as a unique cluster,

etc. etc.

But, I may be wrong.

Do you suppose that the distribution plotted on Bruce's chart could

indicate the authorship of a single writer (minus Isaiah/Malachi)?

If so, I'd like to hear what you have to say about that expansive pattern.

UD

Posted

It seems reasonable that one could use the Moroni chapters (except #9), to part out what Spaldingites know is not Spalding. And Biblical chapters, and distinctly Smithian chapters, to separate out what all agree is not Spalding. Not necessarily in that order. Theoretically, what is left would be Spalding?

Posted

If I knew the answer to that, I'd have used it already. I'm not sure there is a perfect methodology for this purpose, but then I'm not an expert in statistical authorship attribution.

Peace,

-Chris

:P Excellent work, Chris. I am on board with you here. Especially with a reticence to accept wordprint studies as very reliable.

Posted

No confusion at all now, Bruce. You've articulated all the main points, I think.

I assume that if the process were reversed (charting closest non-variations from

the mean), that we'd simply end up with a huge cluster around the 0,0 point?

Are you asking whether we could find the 2 linear combinations of features that display the least possible spread? Sure, but if you were to plot the points in this 2-dimensional space using default settings of plotting software, the points would likely be spread randomly all over the page, with no clustering pattern of any kind (by default the plotting scales are chosen such that the plot covers as much of the plotting space as possible).

If you've followed the ongoing discussion here, over a considerable period

of time, in many threads, you'll recall that the vocal advocates for sundry

BoM origin explanations fall into three general categories:

1. Jaredites and Nephites wrote it; Smith translated into his vernacular

2. Smith wrote it by himself (incorporating a little KJV material)

3. Smith also incorporated some non-biblical pre-existing texts

I'm not sure just how much your chart can help us here -- to better

define and "rate" those three possibilities -- but you can imagine my

interest in the question.

I

Archived

This topic is now archived and is closed to further replies.

  • Recently Browsing   0 members

    • No registered users viewing this page.
×
×
  • Create New...