In order to diagnose and treat knee osteoporosis effectively, it is very important for doctors to be able to do both quickly and accurately. In this paper we discuss a way to create a deep-learning hybrid model that will combine Convolutional Neural Network (CNN) algorithms with Multimodal FineTuned Large Language Models (LLM) in an effort to create an improved diagnostic tool. Traditional CNN architectures such as InceptionV3, DenseNet121, MobileNetV2 and ResNet50 were utilized in our study and found to have moderate accuracy of about 80%. This is not high enough for a tool to be relied upon in clinical settings. Subsequently a Multimodal Methodology was introduced which yielded minimal performance improvements but again did not meet our expectations. An advanced Multimodal LLM and CNN combination model was developed and allowed us to take advantage of the complex relationships that are inherent in images and their associated meta-data. The final results were a diagnostic accuracy of 93%. This indicates that our combination method produced a significantly higher accuracy level than either traditional CNNbased diagnostic tools or previous Multimodal methodologies. Therefore, by combining the advantages of both CNNs and Additionally Fine-Tuned LLMs, the direction is encouraging for improving the diagnosis of osteoporosis via diagnostic imaging.