What is PySpark and how is it used in big data processing and data engineering? How does PySpark enable distributed data processing using Apache Spark? What are the key features and components of PySpark? How is PySpark used for data cleaning, transformation, and machine learning tasks? What are the advantages and limitations of using PySpark in real-world data workflows?