فصلنامه سیستم های انرژی پایدار

فصلنامه سیستم های انرژی پایدار

بهینه‌سازی دینامیکی سیستم تولید همزمان مبتنی بر سیکل کالینا با استفاده از یادگیری تقویتی عمیق

نوع مقاله : مقاله پژوهشی

نویسندگان
1 دانشکده مهندسی انرژی و منابع پایدار، دانشکدگان علوم و فناوریهای میان رشته ای، دانشگاه تهران، تهران، ایران
2 دانشکده مهندسی انرژی و منابع پایدار، دانشکدگان علوم و فناوری های میان رشته ای دانشگاه تهران، تهران، ایران
3 دانشکده مهندسی انرژی و منابع پایدار، دانشکدگان علوم و فناوری های میان رشته ای دانشگاه تهران، تهران، ایران.
10.22059/ses.2026.415139.1248
چکیده
در این پژوهش، یک چارچوب یکپارچه برای مدل‌سازی ترمودینامیکی-اقتصادی و بهینه‌سازی پویای بهره‌برداری از یک سیستم تولید همزمان برق، گرمایش و سرمایش مبتنی بر سیکل کالینا و چیلر جذبی ارائه شده است. مسئله بهره‌برداری در یک افق زمانی ۲۴ ساعته با گام‌های یک‌ساعته صورت‌بندی و الگوریتم یادگیری تقویتی عمیق Soft Actor–Critic (SAC) برای تصمیم‌گیری پیوسته به کار گرفته شد. سه سیاست کنترلی سودمحور، متعادل و بازده‌محور طراحی و عملکرد آن‌ها بر اساس پنج اجرای مستقل با سیاست مرجع مبتنی بر برنامه‌ریزی خطی مقایسه شد. نتایج نشان داد که سود عملیاتی روزانه سیاست مرجع برابر با 419.03 یورو و شاخص عملکرد انرژی آن 0.07724 است. میانگین سود روزانه سیاست‌های سودمحور، متعادل و بازده‌محور به‌ترتیب 538.318، 547.701 و 589.236 یورو و شاخص عملکرد انرژی آن‌ها به‌ترتیب 0.08036، 0.08014 و 0.08099 به دست آمد. این نتایج معادل افزایش سود به‌ترتیب 28.5، 30.7 و 40.6 درصد و بهبود شاخص عملکرد انرژی به‌ترتیب 4.0، 3.8 و 4.8 درصد نسبت به سیاست مرجع است. سیاست بازده‌محور بالاترین میانگین هر دو شاخص را ارائه کرد، هرچند پراکندگی نتایج اجراهای مستقل، حساسیت عملکرد به تصادفی‌بودن فرایند آموزش را نشان داد. همچنین، زمان استنتاج کمتر از یک میلی‌ثانیه، قابلیت بالقوه استفاده برخط از سیاست آموزش‌دیده را از نظر محاسباتی نشان داد. تحلیل حساسیت نیز بیانگر تأثیر مستقیم تغییرات قیمت برق بر عملکرد اقتصادی و حساسیت محدود خروجی‌های ارزیابی‌شده سیاست یادگیری تقویتی نسبت به تغییرات تقاضا در محدوده بررسی‌شده بود.‌
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Dynamic Optimization of a Kalina-Based CCHP System Using Deep Reinforcement Learning

نویسندگان English

Ali Roghani Araghi 1
Hossein Yousefi 2
Saba Amouzade 3
Mona Mirrazavi 2
1 School of Energy Engineering and Sustainable Resources, College of Interdisciplinary Science and Technology, University of Tehran, Tehran, Iran
2 , Faculty of Energy Engineering and Sustainable Resources, College of Interdisciplinary Sciences and Technologies, University of Tehran, Tehran, Iran
3 Faculty of Energy Engineering and Sustainable Resources, College of Interdisciplinary Sciences and Technologies, University of Tehran, Tehran, Iran
چکیده English

This study presents an integrated framework for thermodynamic–economic modeling and dynamic operational optimization of a combined cooling, heating, and power system based on a Kalina cycle and an absorption chiller. The operational problem was formulated over a 24-hour horizon with hourly time steps, and the deep reinforcement learning Soft Actor–Critic (SAC) algorithm was employed for continuous decision-making. Three control policies, namely profit-oriented, balanced, and efficiency-oriented policies, were developed, and their performance, based on five independent runs, was compared with a linear programming-based reference policy. The results showed that the reference policy achieved a daily operating profit of €419.03 and an energy performance index of 0.07724. The mean daily profits of the profit-oriented, balanced, and efficiency-oriented policies were €538.318, €547.701, and €589.236, respectively, while their corresponding energy performance indices were 0.08036, 0.08014, and 0.08099. These results correspond to profit improvements of 28.5%, 30.7%, and 40.6% and energy performance index improvements of 4.0%, 3.8%, and 4.8%, respectively, compared with the reference policy. The efficiency-oriented policy achieved the highest mean values for both performance indicators, although the dispersion across independent runs demonstrated the sensitivity of performance to the stochastic nature of the training process. Moreover, an inference time of less than one millisecond indicated the computational potential of the trained policy for online implementation. Sensitivity analysis further demonstrated the direct effect of electricity price variations on economic performance and the limited sensitivity of the evaluated reinforcement learning policy outputs to demand variations within the investigated range.

کلیدواژه‌ها English

Combined Cooling Heating and Power (CCHP)
Kalina Cycle
Deep Reinforcement Learning
Dynamic Optimization
Thermodynamic-Economic Modeling

مقالات آماده انتشار، پذیرفته شده
انتشار آنلاین از 31 تیر 1405