Textual analysis of Book of Mormon timeline proves it could not have been faked

This research is very similar to Statistical Analysis of 368 Accounts of the Flood Proves both the Book of Mormon and the Bible are True, so it will probably make more sense if you read that first.

The most important finding here is that if one tries to predict the position of a verse in the Book of Mormon using linear regression and out of sample predictions, the results are much more correlated with the time period claimed in the Book of Mormon chapters than they are with the actual position of the verse. In fact, using OLS, you actually cannot predict the position of a verse in the Book of Mormon from my dataset (which I will explain shortly) – you *only* can predict the time period the Book of Mormon claims the verse was written.

If the Book of Mormon were faked, one would expect the opposite to be true: the natural progression of ideas and textual styles found in the verses of the Book of Mormon should be more correlated with the position of the verse in the Book of Mormon than the fake dates provided in the Book of Mormon for each chapter. On the other hand, if the Book of Mormon actually is true, this result is exactly what you would expect: since the Book of Mormon is said to be comprised of various writings of prophets across a thousand year span, you would expect to be able to predict when these writings were, since each prophet would have a writing signature typical for his time period.

These results align with work done by others, discussed in Secular Evidence for the Book of Mormon, statically showing that there are multiple authors to the Book of Mormon – as opposed to just one author, Joseph Smith. Regarding these multiple authors, anti-Mormons will say that this could have been faked by a clever Joseph Smith – but research shows that this would have been nearly impossible, given that the statistical score for diversity in writers in the Book of Mormon is actually much higher than the top contemporaneous fiction writers in Joseph Smith’s time – and even higher than several of these writer’s works all combined together. Or in other words, if the top authors in Joseph Smith’s time period – all combined together – could not produce as diverse a set of voices as the Book of Mormon, it is unrealistic to think that Joseph Smith – an uneducated 23 year old – could have done so.

And just to emphasize one part of my present research here: when I try to predict the position of a verse in the Book of Mormon, the prediction ends up being much more correlated with the date claimed in the Book of Mormon of that verse than it is with the actual position of the verse. Additionally, the predictions for the actual position of the verse, using out of sample predictions, are so bad that they are not statistically significant. Or in other words: you cannot predict, using my methods which analyze the textual signature of each verse, the position of a verse in the Book of Mormon. But you can easily predict the time period in which the Book of Mormon claims each verse was written.

However, if the model instead tries to predict the claimed time period of each verse, the resultant predictions actually are correlated with the position of the verse. Therefore, if you directly try to predict verse position, using the actual verse position as the y variable in training – then this doesn’t work. But if you instead indirectly try to predict the verse position by using the claimed time period of each chapter in the Book of Mormon, then this works just fine. This again aligns with the Book of Mormon – because since the Book of Mormon is written in a chronologically strange way – where some time periods are covered very extensively, whereas others are covered briefly – you would expect models that are trained on verse position to be kind of weird and messed up, but models trained on date to be OK.

But that’s just a summary. Now I will fully present the methods used to produce these results and further detail the results.

Data

To conduct this analyses, I ran each verse of the Book of Mormon into the AI model llama3.1:8b, asking the model to provide a score for each verse across 264 fields. This gave me a 264 item vector representing the contents of every verse.

Applying clustering analysis to the resultant dataset, here is a link to a figure of the clusters of all 264 fields – showing which fields tend to be more correlated with each other.

Field Clusters

Methods

After decomposing the text into vectors, I merged this dataset into another dataset of average years for each Book of Mormon chapter, giving a rough estimate of the year the Book of Mormon claims each verse was written.

Since 264 fields is far too many fields to predict the year and position of each verse, I then took the first 20 principal components of the 264 fields, and conducted my analysis on these components. In brief, principal components are new artificial variables created to combine features in a way that explains as much of the variance of the original features as possible. This is overwhelmingly the most common way of simplifying data when there are concerns with overfitting.

To actually create the predictions for the year of each verse, I went through each sub-book in the Book of Mormon (the Book of Mormon contains several different sub-books, such as 1 Nephi, Alma, Ether, and several others). I then trained an OLS regression model on the data from all of the Book of Mormon verses *not* in the particular sub-book in question, and then used the model to predict the year of each verse within said sub-book. Repeating this process for each sub-book, this allowed me to form predictions for the year of each verse, for which the specific model creating the predictions never actually was trained on that particular verse. This is called leave one group out cross validation.

Since I leave each sub-book out for the predictions, you, the skeptic, shouldn’t even try to tie your head in a knot saying that my results are wrong because I don’t take into account within sub-book textual similarities – because my results for each sub-book are never found using data from that book.

Using this method, I was able to create predictions for 1) the year each verse was said to be written, 2) the position of that verse in the Book of Mormon (for example, one verse might be the 6298th verse in the Book of Mormon), and 3) the position of the sub-book of that verse in the Book of Mormon (for example, 2-Nephi is the second sub-book in the Book of Mormon, and Omni is the sixth sub-book in the Book of Mormon.

Results

Here are the correlations and p-values for the predicted values and the actual values. For context, a p-value is the chance that the correlation could happen by chance – so if a p-value is low, then the correlation significant.

Let me go through these results one by one:

Starting with the green rows, when I predict bookID, that means I am trying to predict the position of the sub-book that the verse is in. Therefore, if the Book of Mormon were a fraud, I would expect this prediction to be more correlated with the position of the verse in the larger Book of Mormon, then with the year the Book of Mormon claims the verse was written. But we see the opposite: this prediction is actually more correlated with the year the verse was claimed to be written than it is with either the verse position or the actual position of the sub-Book in the Book of Mormon. This is so extreme that the correlation with verse-position is actually slightly negative.

Moving to the blue rows, these are trained to predict the position of the particular verse in the Book of Mormon. As you can see, the predictions just completely fail – however, this logic is actually able to predict the claimed year the verse was written. I think this is because the year the chapter is said to have been published will be correlated with the verse position, since the Book of Mormon is mostly written in chronological order. But since the verse position is contrived (since the Book of Mormon somewhat arbitrarily chooses to highly select from certain years while ignoring others in its selections from the writings of the prophets), any sort of model trying to predict verse position will do poorly, but since the year the verse was written is more real (since there are distinct writing styles for each year), these sorts of models will work better when predicting the year the verse was published.

Finally, on the purple rows, we can see that the predicted year the verse was published is able to predict all three metrics – the year, the position of the verse, and the position of the sub-book. Oddly, its predictions are slightly better when predicting the verse position than the claimed year, but this difference is small and doesn’t negate my results (and as you will see in the next section, this oddity goes away). The more important thing to notice is that if you want to predict *the position a verse was published*, models that are trained on the year the verse is claimed to have been written actually work, whereas models trained on the position of the sub-verse don’t work. This is because, again, from a textual analysis perspective, it is difficult to tell where a verse is placed in the Book of Mormon, since this position is arbitrarily hugely distorted when the Book of Mormon emphasizes certain years over others – but it is much easier to do this looking at the actual date of each verse, because the dates are associated with real textual styles from the time period that can be seen statistically.

And to note – if you look at the p-values, my results are very, very, significant. For example, the p-value that the predicted verse position is correlated with the claimed year the verse was written is 2.12 times ten to the negative 83.

Sanity Check – Excluding Ether

One oddity in the Book of Mormon is that while it is mostly in chronological order, the second to the last book – the book of Ether – actually comes from a prior date to all of the other books in the Book of Mormon.

Therefore, one could imagine that my results simply are due to some oddness in the Book of Ether, and if I exclude Ether, then my results will no longer be significant. So these are the results if I do the same calculations as before, but excluding Ether.

As you can see, all of the results are the same: 1) predicted sub-book position is more correlated with claimed year than verse position, 2) the predicted verse position is still way more correlated with claimed year than actual verse position, and 3) models trained on claimed year perform better at predicting verse position than models trained on verse position.

Conclusions

This research shows that the statistical signatures of verses in the Book of Mormon are vastly and immensely more corelated with the year the Book of Mormon claims each verse was written than they are with the actual position of the verse in the Book of Mormon. This cannot be attributed to specific styles for each sub-book in the Book of Mormon, because this analysis employs leave one group out cross validation for the sub-books. These results are so extreme that when one tries to predict the position of a verse in the Book of Mormon (for example, Mosiah 28:19 is the 2,465th verse in the Book of Mormon), models which are actually trained on the claimed year the verse was written, and not in fact the position of the verse, do immensely better than models directly trained on the position of the verse. Because of this, all I can conclude is that the Book of Mormon must actually be true.

To read the next part of this research, go to The Book of Mormon provides unique insight explaining the statistical structure of how doctrines are presented in the Bible.

Finally, here is the code and data used to produce these statistics:


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Index