|
Arrowsmith Home Word and Text Similarity Metrics |
Word and Text Similarity MetricsWords, bigrams, trigrams and abbreviations. A database of all PubMed articles published 1966-2016, in English (or with titles and abstracts translated into English), was created. Titles and abstracts were tokenized and stoplisted using the PubMed 365-word stoplist. Words were included only if they appeared in at least 100 titles and in at least 25 abstracts. The basic implicit word model consists of 44,201 words that met these criteria and includes 41,918,680 word pairs that co-occurred in the title or abstract of the same PubMed article. For the full model, we included 43,494 words, 49,918 bigrams, 3,638 trigrams, and 704 abbreviations and 92,757,119 co-occurring term pairs. Computing vector representations for terms. For each term selected as above, we made a list of all other selected terms that co-occurred with it in titles and abstracts of MEDLINE articles, and calculated their Direct Odds Ratio. This is a normalized measure of co-occurrence frequency that measures how often two words are observed to co-occur in the same article, relative to the value that would be expected by chance (i.e., if each word occurred at random throughout the corpus of PubMed articles and independently of each other). Then, we ordered the list according to their direct odds ratios. The co-occurring terms with the highest direct odds ratios were used to form a vector. That is, for a given term, we took its co-occurring term having the highest direct odds ratio and placed it in dimension 1; took the term having the next highest direct odds ratio and placed it in dimension 2; and so on, until 300 dimensions were assigned or until the direct odds ratio fell below 1.25. Thus, the vector representation of a word or term consists of a ranked list of its co-occurring context terms. Computing implicit similarity metrics for terms. a) Implicit Shared Terms. For any two terms in our dataset, we examined their 300-dimensional vectors, and counted the number of words (or terms) shared in both vectors. The unweighted similarity is thus an integer ranging from 0 to 300. b) Implicit Weighted Score for the basic (single word) model. For two words or terms A and B, we examined their 300-dimensional vectors, and list the words shared in both vectors. We then created a weighted sum as follows: For each shared vector word, choose the LESSER of the odds ratio in the two vectors. Then, the implicit weighted score = the sum of the loge odds ratios across all shared words. c) Implicit Weighted Score for the full (term) model. We give greater weight for shared trigrams and bigrams relative to shared words or abbreviations. Thus, we create a weighted sum as follows: For each shared vector term, choose the LESSER of the direct odds ratio in the two vectors. Then, the implicit weighted score = the sum of (wi* loge odds ratios) across all shared words, where wi = 1 for words and abbreviations, 2 for bigrams, 3 for trigrams. Computing vector representations of text passages. For each PubMed article, the title+abstract was concatenated to form a single text passage, and tokenized and processed to recognize words, bigrams, trigrams, abbreviations from our dataset. Each term was only counted once, i.e., there was no double-counting for the title or for multiple instances in the abstract. For each term found in the title+abstract, its 300-dimensional vector representation was listed. Then, we created a master list of all terms that occur across all these vectors. Each term was assigned a value which is the sum of the loge direct odds ratio of that term in each vector it appears in. Finally, we represented the PubMed title+abstract as a 300-dimensional vector as follows: The term having the highest value gets rank 1, next highest is rank 2, etc. down to 300 dimensions. Computing implicit similarity metrics for text passages. For two articles (title+abstract) or other text passages, we examine their 300-dimensional vector representations, and listed the words (or terms) shared in both vectors. The implicit shared terms metric is simply the number of shared terms, i.e., an integer between 0 and 300. The implicit weighted score is computed as follows: For each shared vector term, choose the LESSER of the direct odds ratio in the two vectors. Then, the implicit weighted score = the sum of (wi* loge odds ratios) across all shared terms, where wi = 1 for words and abbreviations, 2 for bigrams, 3 for trigrams. Word2vec based similarity metrics. For computing word2vec term similarity, we used the University of Turku word2vec 200-dimensional vectors computed for biomedical words that was trained on PubMed titles and abstracts downloaded from http://evexdb.org/pmresources/vec-space-models/. This vector representation is referred to here as "word2vec". To obtain the word2vec similarity of two words or phrases, the cosine similarity of their word2vec vectors was computed, giving a real number between -1 and +1. Word2vec-based pvtopic vector representation of textual passages and similarity metrics. To represent PubMed articles, word2vec 300-dimensional vectors were computed using the paragraph2vec code https://github.com/hassyGo/paragraph-vector using the following parameters: -wvdim 300 -pvdim 300 -itr 10. Training was performed across the entire PubMed corpus of article titles and abstracts (1966-2016). This training procedure created a separate word2vec representation of biomedical words which will be referred to as "pvtopic" since it follows the word2vec vector representation described in Hashimoto et al. To obtain the similarity of two text passages, the cosine similarity of their vectors were computed. To ensure robustness, similarity scores were only computed for articles that contained abstracts and whose title+abstract contained at least 25 words. Last modified: January 4, 2018 |