What is GPT (Generative Pre-trained Transformer) and how does it work in natural language processing? What are the key components of the transformer architecture used in GPT models? How does pre-training and fine-tuning improve the performance of GPT? What are the common applications of GPT in real-world scenarios? What are the limitations and ethical concerns associated with GPT models?