What is momentum in deep learning and how does it improve the training process of neural networks? How does momentum help accelerate gradient descent optimization? What problems in model training can momentum help solve? How does momentum differ from standard gradient descent methods? What are the advantages and limitations of using momentum in deep learning optimization?