Exploring BERT’s sensitivity to lexical cues using tests from semantic priming

  author={Kanishka Misra and Allyson Ettinger and Julia Taylor Rayz},
Models trained to estimate word probabilities in context have become ubiquitous in natural language processing. How do these models use lexical cues in context to inform their word probabilities? To answer this question, we present a case study analyzing the pre-trained BERT model with tests informed by semantic priming. Using English lexical stimuli that show priming in humans, we find that BERT too shows “priming”, predicting a word with greater probability when the context includes a related… 

