This paper proposes a deep neural network (DNN) based non-intrusive speech quality estimation method in real-time voice communication systems. Since the proposed method only utilizes real-time control protocol (RTCP) information in the receiver side and does not need a reference signal, it is possible to continuously monitor the quality of service (QoS). Unlike the conventional non-intrusive E-model system that predicts QoS by utilizing delay, jitter, and type of codec with a rule-based method, the proposed method actively estimates the non-linear relationship between multi-dimensional parameters of RTCP and subjectively motivated reference scores using a DNN structure. In order to select efficient features, the relationship between each parameter of RTCP and perceptual objective listening quality assessment (POLQA) is thoroughly investigated, then we train the DNN model by changing the number of layers and nodes. The proposed algorithm achieved 0.8693 correlation with 21,206 reference POLQA scores that are sampled from real environment.