Tasks that are efficient to reduce latency and make the best use of resources in distributed networks, it is important to program in fog computing. Conventional programming algorithms frequently struggle to accommodate dynamic workloads and diverse resources. In this research, we introduce DQS (Deep Q Learning Planner), a framework grounded in reinforcement learning (RL) that utilizes a hybrid Markov Decision Process (POMDP) and an optimized reward function to enhance task success and energy efficiency. DQS is assessed in various reference scenarios and is compared with standard reinforcement learning methodologies, including Q-Learning and Sarsa, as well as traditional heuristic programming techniques. Experimental results indicate that DQS attains a success rate of 9 4. 2 %±0. 9 % for tasks, decreases average latency by 18 %, and enhances energy efficiency by 12 % relative to the optimal performance baseline. DQS also has ways to handle failures to make sure it works reliably even when the network is not stable. These results show that combining a hybrid state representation with a carefully designed reward function makes programming fog tasks much easier. This is a strong, flexible, and scalable solution that works well in real-world computing environments.