Circular Image

M.A. Sharifi Kolarijani

info

Please Note

6 records found

Journal article (2024) - Yichen Liu, Mohamad Amin Sharifi Kolarijani
In this letter, we consider the application of max-plus-linear approximators for Q-function in offline reinforcement learning of discounted Markov decision processes. In particular, we incorporate these approximators to propose novel fitted Q-iteration (FQI) algorithms with provable convergence. Exploiting the compatibility of the Bellman operator with max-plus operations, we show that the max-plus-linear regression within each iteration of the proposed FQI algorithm reduces to simple max-plus matrix-vector multiplications. We also consider the variational implementation of the proposed algorithm which leads to a per-iteration complexity that is independent of the number of samples. ...
Journal article (2023) - M.A. Sharifi Kolarijani, P. M. Esfahani
We propose two novel numerical schemes for the approximate implementation of the dynamic programming (DP) operation concerned with finite-horizon optimal control of discrete-time systems with input-affine dynamics. The proposed algorithms involve discretization of the state and input spaces and are based on an alternative path that solves the dual problem corresponding to the DP operation. We provide error bounds for the proposed algorithms, along with a detailed analysis of their computational complexity. In particular, for a specific class of problems with separable data in the state and input variables, the proposed approach can reduce the typical time complexity of the DP operation from O(XU) to O(X+U) , where X and U denote the size of the discrete state and input spaces, respectively. This reduction in complexity is achieved by an algorithmic transformation of the minimization in DP operation to an addition via discrete conjugation. ...

With Applications in Opinion Dynamics and Dynamic Programming

This thesis is comprised of two main parts. In the first part of the thesis, we study the nonlinear Fokker-Planck (FP) equation that arises as a mean-field (macroscopic) approximation of the bounded confidence opinion dynamics, where opinions are influenced by environmental noises and opinions of radicals (stubborn individuals). The distribution of radical opinions serves as an infinite-dimensional exogenous input to the FP equation, visibly influencing the steady opinion profile. We first establish the mathematical properties of the FP equation. In particular, we (i) show the well-posedness of the dynamic equation, (ii) provide existence result accompanied by a quantitative global estimate for the corresponding stationary solution, and, (iii) establish an explicit lower bound on the noise level that guarantees exponential convergence of the dynamics to stationary state. Combining the results in (ii) and (iii) readily yields the input-output stability of the system for sufficiently large noises. Next, using Fourier analysis, the structure of opinion clusters under the uniform initial distribution is examined. Specifically, two numerical schemes for (i) identification of order-disorder transition and (ii) characterization of initial clustering behavior are provided. The results of the analysis are validated through several numerical simulations of the continuum-agent model (partial differential equation) and the corresponding discrete-agent model (interacting stochastic differential equations) for a particular distribution of radicals.

In the second part of the thesis, we focus on the value iteration algorithm for solving optimal control problems. We propose two novel numerical schemes for approximate implementation of the dynamic programming (DP) operation concerned with finite-horizon, optimal control of deterministic, discrete-time systems with input-affine dynamics. The proposed algorithms involve discretization of the state and input spaces and are based on an alternative path that solves the dual problem corresponding to the DP operation. We provide error bounds for the proposed algorithms, along with detailed analyses of their computational complexity. In particular, for a specific class of problems with separable data in the state and input variables, the proposed approach can reduce the typical time complexity of the DP operation from O(XU) to O(X +U), where X and U denote the size of the discrete state and input spaces, respectively. We next discuss the extensions of the proposed conjugate value iteration algorithm for problems with separable data. The extensions are three-fold: We consider (i) infinite-horizon, discounted cost problems with (ii) stochastic dynamics, while (iii) computing the conjugate of input cost numerically. In particular, we analyze the convergence, complexity, and error
of the proposed algorithm under these extensions. The theoretical results are validated through multiple numerical examples. ...
In this study, we consider the infinite-horizon, discounted cost, optimal control of stochastic nonlinear systems with separable cost and constraints in the state and input variables. Using the linear-time Legendre transform, we propose a novel numerical scheme for implementation of the corresponding value iteration (VI) algorithm in the conjugate domain. Detailed analyses of the convergence, time complexity, and error of the proposed algorithm are provided. In particular, with a discretization of size X and U for the state and input spaces, respectively, the proposed approach reduces the time complexity of each iteration in the VI algorithm from O(XU) to O(X+U), by replacing the minimization operation in the primal domain with a simple addition in the conjugate domain. ...
Journal article (2021) - Mohamad Amin Sharifi Kolarijani, Anton V. Proskurnikov, Peyman Mohajerin Esfahani
In this article, we study the nonlinear Fokker-Planck (FP) equation that arises as a mean-field (macroscopic) approximation of bounded confidence opinion dynamics, where opinions are influenced by environmental noises and opinions of radicals (stubborn individuals). The distribution of radical opinions serves as an infinite-dimensional exogenous input to the FP equation, visibly influencing the steady opinion profile. We establish mathematical properties of the FP equation. In particular, we, first, show the well-posedness of the dynamic equation, second, provide existence result accompanied by a quantitative global estimate for the corresponding stationary solution, and, third, establish an explicit lower bound on the noise level that guarantees exponential convergence of the dynamics to stationary state. Combining the results in second and third readily yields the input-output stability of the system for sufficiently large noises. Next, using Fourier analysis, the structure of opinion clusters under the uniform initial distribution is examined. The results of analysis are validated through several numerical simulations of the continuum-agent model (partial differential equation) and the corresponding discrete-agent model (interacting stochastic differential equations) for a particular distribution of radicals. ...
Conference paper (2019) - M. A.S. Kolarijani, A. V. Proskurnikov, P. Mohajerin Esfahani
In this paper, we consider the mean-field model of noisy bounded confidence opinion dynamics under exogenous influence of static radical opinions. The long-term behavior of the model is analyzed by providing a sufficient condition for exponential convergence of the dynamics to stationary state. The stationary state is also characterized by a global estimate for a sufficiently large noise. Furthermore, we consider the order-disorder transition in the model in order to identify the effect of the (relative) mass of the radicals on the critical noise level at which this transition occurs. A numerical scheme for approximating the critical noise level is provided and validated through numerical simulations of the mean-field model and the corresponding agent-based model for a particular distribution of radical opinions. ...