What is weight initialization in deep learning and why is it important for training neural networks? How do Xavier and He initialization methods work in setting initial weights? How does proper weight initialization improve training stability and convergence? What is the difference between Xavier initialization and He initialization? What are the advantages and limitations of different weight initialization techniques?