Arrowsmith Home

Word and Text Similarity Metrics

In this webpage, we present several different ways of representing words, multi-word phrases, abbreviations, and textual passages as low-dimensional vectors, and employ the vectors to create several different ways of computing term-term and text-text similarity metrics. Some of these metrics are new (devised by us) and some are derived from previous research by others. PubMed titles and abstracts have been employed as a corpus, and we hope that the resulting datasets and query interfaces should be valuable both for biomedical investigators and for text mining groups. Click HERE for a brief summary of the vector representations and similarity metrics. See the accompanying reference (below) for full details.

Term Vector Representations and Similarity Metrics

Basic Model Full Model (includes bigrams, trigrams and abbreviations)
Search Word Similarity Metrics


Text Vector Representations and Similarity Metrics

Implicit vectors based on the full model have been computed for the entire PubMed corpus of articles (1966-2016) published in English (or with titles and abstracts translated into English) and whose title+abstract contained a minimum of 25 words. The data is available for download at:
Compute similarity of two PubMed title+abstracts according to Implicit Weighted Score

Pvtopic vectors have been computed for the entire PubMed corpus of articles (1966-2016) that contained abstracts and whose title+abstract contained a minimum of 25 words, and the data can be downloaded using the links below.
Compute similarity of two PubMed title+abstracts according to Pvtopic Cosine Similarity

Updated Pvtopic vectors have been computed for the entire PubMed corpus of articles (1966-2020) whose title+abstract contained a minimum of 25 words. New vectors have been computed with the gensim python library to allow vectors for new articles to be inferred without remodeling the entire corpus. The data can be downloaded using the links below.
Reference

Smalheiser NR, Bonifield G. (2018) Unsupervised Low-Dimensional Vector Representations for Words, Phrases and Text that are Transparent, Scalable, and produce Similarity Metrics that are Complementary to Neural Embeddings. Preprint deposited in arXiv on January 9, 2018. http://arxiv.org/abs/1801.01884

Supplemental files.



These data are being released under the terms of the Creative Commons Attribution-NonCommercial-ShareAlike CC BY-NC-SA International Public License 4.0.

Last modified: January 15, 2018