Radar-based human detection has gained significant importance in surveillance, security, and smart monitoring systems due to its robustness in low-visibility conditions and privacy-sensitive environments. While recent studies have explored deep learning for radar data analysis, they typically specialize in either one-dimensional time-series or two-dimensional image representations, potentially limiting their feature extraction capability. This study proposes a unified multimodal convolutional neural network framework that simultaneously processes both 1D time-series radar signals and 2D radar image representations from the same Ultra-Wideband radar acquisitions. The approach addresses critical data challenges including variable-length sequences through zero-padding and significant class imbalance through dataset stratification. Experimental results on a dataset of 283 samples demonstrate that the proposed model achieves 91.2% accuracy with balanced precision, recall, and F1-scores across both human and non-human classes. Despite dataset constraints, the multimodal approach demonstrates an approximately 8% accuracy improvement over single-modality baselines, validating the feasibility of complementary feature fusion. This work establishes a foundational framework for multimodal radar processing and provides insights for future development of human detection systems leveraging multiple data representations.