维普中文期刊产品整合服务
6篇 您的检索式:作者名="Abhijit GOSAVI"
    题名 作者 年代 出处 被引量
1A reinforcement learning approach to a single leg airline revenue management problem with multiple fare classes and overbooking显示文摘Abhijit Gosavi Naveen Bandla Tapas K. Das 2002IIE Transactions2002,,9:1
2Solving Markov Decision Processes with Downside Risk Adjustment显示文摘Markov decision processes(MDPs) and their variants are widely studied in the theory of controls for stochastic discreteevent systems driven by Markov chains.Much of the literature focusses on the risk-neutral criterion in which the expected rewards,either average or discounted,are maximized.There exists some literature on MDPs that takes risks into account.Much of this addresses the exponential utility(EU) function and mechanisms to penalize different forms of variance of the rewards.EU functions have some numerical deficiencies,while variance measures variability both above and below the mean rewards;the variability above mean rewards is usually beneficial and should not be penalized/avoided.As such,risk metrics that account for pre-specified targets(thresholds) for rewards have been considered in the literature,where the goal is to penalize the risks of revenues falling below those targets.Existing work on MDPs that takes targets into account seeks to minimize risks of this nature.Minimizing risks can lead to poor solutions where the risk is zero or near zero,but the average rewards are also rather low.In this paper,hence,we study a risk-averse criterion,in particular the so-called downside risk,which equals the probability of the revenues falling below a given target,where,in contrast to minimizing such risks,we only reduce this risk at the cost of slightly lowered average rewards.A solution where the risk is low and the average reward is quite high,although not at its maximum attainable value,is very attractive in practice.To be more specific,in our formulation,the objective function is the expected value of the rewards minus a scalar times the downside risk.In this setting,we analyze the infinite horizon MDP,the finite horizon MDP,and the infinite horizon semi-MDP(SMDP).We develop dynamic programming and reinforcement learning algorithms for the finite and infinite horizon.The algorithms are tested in numerical studies and show encouraging performance.Abhijit Gosavi Anish Parulekar 2016International Journal of Automation and computing2016,13,3:1
3A Reinforcement Learning Approach to a Single Leg Airline Revenue Management Problem with Multiple Rare Classes and Overbooking 显示文摘Abhijit Gosavi Naveen Bandla Tapas K Das 2002IIE Transactions2002,34,9:1
4Semi-Markov adaptive critic heuristics with application to airline revenue management显示文摘The adaptive critic heuristic has been a popular algorithm in reinforcement learning(RL) and approximate dynamic programming(ADP) alike.It is one of the first RL and ADP algorithms.RL and ADP algorithms are particularly useful for solving Markov decision processes(MDPs) that suffer from the curses of dimensionality and modeling.Many real-world problems,however,tend to be semi-Markov decision processes(SMDPs) in which the time spent in each transition of the underlying Markov chains is itself a random variable.Unfortunately for the average reward case,unlike the discounted reward case,the MDP does not have an easy extension to the SMDP.Examples of SMDPs can be found in the area of supply chain management,maintenance management,and airline revenue management.In this paper,we propose an adaptive critic heuristic for the SMDP under the long-run average reward criterion.We present the convergence analysis of the algorithm which shows that under certain mild conditions,which can be ensured within a simulator,the algorithm converges to an optimal solution with probability 1.We test the algorithm extensively on a problem of airline revenue management in which the manager has to set prices for airline tickets over the booking horizon.The problem has a large scale,suffering from the curse of dimensionality,and hence it is difficult to solve it via classical methods of dynamic programming.Our numerical results are encouraging and show that the algorithm outperforms an existing heuristic used widely in the airline industry.Ketaki KULKARNI Abhijit GOSAVI Susan MURRAY Katie GRANTHAM 2011控制理论与应用(英文版)2011,9,3:1
5A Reinforcement Learning Algorithm Based on Policy Iteration for Average Reward: Empirical Results with Yield Management and Convergence Analysis显示文摘Abhijit Gosavi 2004Machine Learning2004,,1:1
6A semi-Markov model for post-earthquake emergency response in a smart city显示文摘在与高人口密度发生在一个区域的 Richter 规模上重要的地震要求一个有效、公平的紧急情况反应计划。紧急情况资源通常位于所谓的反应中心。第一个问题之一在很快降级由灾难反应管理人员面对了地震以后的条件是影响灾难的区域受到的危险率到抵押品,估计花在控制下面带状况的时间,也叫的恢复为 relief-and-rescue 活动预定,并且选择适当反应中心。在这篇论文,我们建议一个精致的 semi-Markov 模型捕获跟随地震的事件的随机的动力学,它将被用来确定危险率人们被暴露并且估计恢复时间到。模型将进一步被使用,经由动态编程,到决定适当反应中心。我们的建议模特儿能与许多危险规模一起并且由在与紧急情况管理有关的一些参数上收集数据被雇用。模型将在一个聪明的城市里是特别地有用的,在跟随地震的事件上的历史性的数据将系统地并且精确地被记录的地方。Shuva GHOSH Abhijit GOSAVI 2017Control Theory and Technology2017,15,1:0
返回顶部 每页显示:
共1页 首页 上一页 第1页 下一页 末页 /1 跳转

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费