pvtopic_2020.tsv is an updated version of the pvtopic vectors described in [1] by N Smalheiser, et al. The model and all inferred vectors were computed with the GenSIM package for Python [2]. Model specifics: Articles were represented as described in [1] with concantenation of title abstract is at least 25 words. Trained on a 50% corpus sample Distributed Memory model with no word down-sampling, no vector concatenation, and using hierarchical softmax optimization 300-dimensional vectors for paragraphs and words 10 “iterations” The exact GenSIM parameters used to fit the model are: model = Doc2Vec(documents=corpus_iterator,dm=1,vector_size=300, window=10, min_count=10, epochs=10, workers=16,alpha=0.025, min_alpha=0.0025, hs=1, dm_concat=0, dm_mean=0, sample=0, negative=0) File specifics: pvtopic2020_vectors.tsv is a tab-delimited text file with two elements. PMID list of 300 vector values separated by spaces There are 18239113 rows as of 2020-04-22. pvtopic2020_wordvectors.tsv is a tab-delimited text file with two elements: word list of 300 vector values separated by spaces There are 508166 rows. References: [1] Unsupervised low-dimensional vector representations for words, phrases and text that are transparent, scalable, and produce similarity metrics that are not redundant with neural embeddings N Smalheiser, A Cohen, G Bonifield - Journal of Biomedical Informatics [2] Software framework for topic modelling with large corpora R Rehurek, P Sojka - In Proceedings of the LREC 2010 Workshop on New... , 2010