BanglaVoice is a curated sentence-level dataset for grammatical voice analysis in Bangla, comprising 4397 annotated sentences categorized as Active (1459), Passive (1531), and Middle (1407). Each instance consists of a structurally complete clause containing a single dominant finite verb, accompanied by its English translation and a categorical voice label. Sentences were selectively compiled from publicly accessible Bangla digital sources published between 2023 and 2025 and underwent systematic cleaning, de-duplication, orthographic normalization, and UTF-8 standardization. Voice annotations were assigned using linguistically defined criteria and validated through multi-annotator agreement with majority voting. The dataset exhibits balanced class distribution and natural language characteristics, including a Zipfian rank–frequency distribution (s ≈ 0.98–0.99; R² ≈ 0.99) and substantial lexical diversity (3234 unique tokens). Baseline experiments using six supervised classifiers are also provided, with LinearSVC achieving 93.18% accuracy and 93.10% F1-score, establishing reproducible reference benchmarks for future Bangla grammatical voice classification research. BanglaVoice is released as an open-access resource to support morpho-syntactic research and voice-aware modelling in Bangla natural language processing.