Home / Current Issue / Paper 1722214
Reinforcement Learning for Intelligent Vehicle-to-Grid Integration: Balancing Clean Energy Utilization and Battery Degradation Preservation
Subject area: Science,Engineering and Technology · Area of research: Reinforcement Learning, Smart Grid, V2G, ML
Abstract
The transition toward sustainable mobility requires an intelligent integration of Electric Vehicles (EVs) into the power grid. This paper proposes a Reinforcement Learning (RL) framework using Proximal Policy Optimization (PPO) to manage bidirectional Vehicle-to-Grid (V2G) power flow. Utilizing the Indian Grid Master dataset, the system optimizes for economic arbitrage and carbon reduction while strictly adhering to a 90% State of Charge (SoC) mobility requirement. A core innovation of this work is an asymmetric reward function that applies a 35x penalty to battery discharge relative to charging, ensuring hardware longevity. Results across various 10-hour shift profiles demonstrate the agent's ability to achieve mobility targets while maximizing grid stability.
Keywords
V2G, Reinforcement Learning, PPO, Battery Degradation, Indian Grid Master, EV2Gym.
How to cite this paper
@article{1722214,
author = {Vansh Sharma, Hrishita Sarkar, Lopamudra Mazumder},
title = {Reinforcement Learning for Intelligent Vehicle-to-Grid Integration: Balancing Clean Energy Utilization and Battery Degradation Preservation},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {2},
pages = {853-859},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722214.pdf},
abstract = {The transition toward sustainable mobility requires an intelligent integration of Electric Vehicles (EVs) into the power grid. This paper proposes a Reinforcement Learning (RL) framework using Proximal Policy Optimization (PPO) to manage bidirectional Vehicle-to-Grid (V2G) power flow. Utilizing the Indian Grid Master dataset, the system optimizes for economic arbitrage and carbon reduction while strictly adhering to a 90% State of Charge (SoC) mobility requirement. A core innovation of this work is an asymmetric reward function that applies a 35x penalty to battery discharge relative to charging, ensuring hardware longevity. Results across various 10-hour shift profiles demonstrate the agent's ability to achieve mobility targets while maximizing grid stability.},
keywords = {V2G, Reinforcement Learning, PPO, Battery Degradation, Indian Grid Master, EV2Gym.},
month = {August},
}