What is cosine similarity, and how is it used in Natural Language Processing tasks? How does cosine similarity measure the similarity between two text vectors? What is the mathematical intuition behind cosine similarity? What are the advantages of using cosine similarity over other distance measures? In which NLP applications is cosine similarity commonly used?