|
Arrowsmith Home Word and Text Similarity Metrics |
Word and Text Similarity MetricsSupplemental File 1. The 30 term relatedness benchmark, showing the 30 terms, their mapping to our dataset, and the relatedness ratings of physicians and coders (obtained from http://rxinformatics.umn.edu/SemanticRelatednessResources.htm). Also shown for each term pair are the similarity values computed according to 5 different metrics: The direct odds ratio, the implicit shared terms, implicit weighted score, word2vec and pvtopic. The latter two metrics are computed in a variety of ways (either using the exact terms or the mapped terms, and either using single word vectors or summed vectors). The Spearman rank correlations are shown for all metrics relative to physician and coder relatedness ratings and among themselves. Supplemental File 2. The 101 term relatedness benchmark, showing the terms, mapping to our dataset, and the relatedness ratings of coders (obtained from http://rxinformatics.umn.edu/SemanticRelatednessResources.htm). One term pair failed to map and was excluded, giving a total of 100 word pairs to evaluate. Also shown for each term pair are the similarity values computed according to 4 different metrics: The direct odds ratio, the implicit shared terms, implicit weighted score, and word2vec. Supplemental File 3. The UMNSRS-Similarity benchmark, showing the 501 term pairs that mapped to our dataset, and the similarity ratings of coders (obtained from http://rxinformatics.umn.edu/SemanticRelatednessResources.htm). Also shown for each term pair are the similarity values computed according to 4 different metrics: The direct odds ratio, the implicit shared terms, implicit weighted score, and word2vec. Supplemental File 4. The UMNSRS-Relatedness benchmark, showing the 511 term pairs that mapped to our dataset, and the mean +/- SD similarity ratings of medical students (obtained from http://rxinformatics.umn.edu/SemanticRelatednessResources.htm). Also shown for each term pair are the similarity values computed according to 4 different metrics: The direct odds ratio, the implicit shared terms, implicit weighted score, and word2vec. Supplemental File 5. The modified UMNSRS-Relatedness benchmark, showing the 420 term pairs that mapped to our dataset, and the mean similarity ratings of medical students (obtained from http://rxinformatics.umn.edu/SemanticRelatednessResources.htm). Also shown for each term pair are the similarity values computed according to 4 different metrics: The direct odds ratio, the implicit shared terms, implicit weighted score, and word2vec. Supplemental File 6. Pairs of PubMed sole-authored articles, including 1,000 pairs randomly chosen that are predicted to be written by the same individual (positive set on Sheet 1), and 1,000 pairs that are predicted to be written by different individuals (negative set on Sheet 2). Shown the pairs of PMIDs, and their direct and implicit similarity scores as computed according to six different metrics (see text). Sheet 3 shows the mean and SD values of the similarity scores in each set, and sheet 4 shows the Spearman rank correlation values among the different metrics. Last modified: January 4, 2018 |