Loading...
Search for: hokmi--sadredin
0.035 seconds

    Decentralized MARL in Road Traffic Congestion in the Presence of Non-Stationarity

    , M.Sc. Thesis Sharif University of Technology Hokmi, Sadredin (Author) ; Haeri, Mohammad (Supervisor)
    Abstract
    In this project, the primary objective is to propose a method to increase the convergence rate and accelerate the time it takes for agents to reach their destination in a traffic network. This topic serves as an intersection between two main areas: reinforcement learning and game theory. In this context, agents cooperate to achieve a common goal while remaining unaware of each other's strategies. Therefore, a decentralized reinforcement learning algorithm is proposed and implemented in the presence of non-stationarity. The proposed algorithm is based on introducing temporary and controlled deviations in the regular reinforcement learning mechanism, specifically the Q-learning algorithm used... 

    Increasing convergence rate in decentralized Q-Learning for traffic networks

    , Article 2024 10th International Conference on Control, Instrumentation and Automation, ICCIA 2024 ; 2024 ; 979-833150997-2 (ISBN) Hokmi, S ; Haeri, M ; Sharif University of Technology
    IEEE  2024
    Abstract
    Given the importance of transportation, the speed of delivering goods and services to their destinations, and the need to prevent congestion in urban traffic networks, the use of effective tools to address these issues is necessary. In this paper, we resolve this problem within the framework of multi-Agent reinforcement learning, where traffic networks are mathematically modeled as a congestion game. Decentralized multi-Agent reinforcement learning shows promise for real-world cooperative tasks where agents lack access to global information, such as the actions of others. While independent Q-learning is frequently employed for decentralized training, the simultaneous policy updates by other... 

    Remote monitoring and control of the 2-DoF robotic manipulators over the internet

    , Article Robotica ; Volume 40, Issue 12 , 2022 , Pages 4475-4497 ; 02635747 (ISSN) Hokmi, S ; Haghi, S ; Farhadi, A ; Sharif University of Technology
    Cambridge University Press  2022
    Abstract
    This article is concerned with remote monitoring and control of the 2-degrees of freedom (DoF) robotic manipulators, which have nonlinear dynamics over the packet erasure channel, which is an abstract model for communication over the Internet, WiFi, or Zigbee modules. This type of communication is subject to imperfections, such as random packet dropout and rate distortion. These imperfections cause a significant challenge for monitoring and control of robotic manipulators in the industrial environments because sensitive data, such as sensor data and control commands may not ever reach to their destination resulting in significant performance degradation. Therefore, the effects of these...