{"id":119,"date":"2020-08-28T12:15:45","date_gmt":"2020-08-28T12:15:45","guid":{"rendered":"https:\/\/eslweb.epfl.ch\/?page_id=119"},"modified":"2020-11-19T08:44:20","modified_gmt":"2020-11-19T08:44:20","slug":"current-master-projects","status":"publish","type":"page","link":"https:\/\/eslweb.epfl.ch\/?page_id=119","title":{"rendered":"Current Master Projects"},"content":{"rendered":"\n<p class=\"has-text-align-right wp-block-paragraph\"><a href=\"https:\/\/eslweb.epfl.ch\/cgi-bin\/projects\/selectprojects.pl\">Edit Master Projects<\/a><\/p>\n\n\n<script>\nfunction opendesc(target,bullet) {\n document.getElementById(target).innerHTML = bullet;\n document.getElementById(target).innerHTML += \"<a href=#_ onClick=closedesc(\\\"\"+target+\"\\\",\"+target+\"minus);>[close]<\/a>\";\n}\n\nfunction closedesc(target,bullet) {\n document.getElementById(target).innerHTML = bullet;\n document.getElementById(target).innerHTML += \"<a href=#_ onClick=opendesc(\\\"\"+target+\"\\\",\"+target+\"); style='position:relative;z-index:99;'> [read&nbsp;on]<\/a>\";\n}\n<\/script>\n\n<div class='container'><div class='row entry-body'><div class='col-sm-4'><br><table width='100%' cellspacing='0' cellpadding='8' border='0' id='table3'><tr><td colspan=2><h3>Master Student Assistant Projects<\/h3><h5>(Remunerated, for officially registered EPFL students only)<\/h5><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor769><\/a><b><span style='font-size: 20px;'>Streaming Wake-Up Networks for Low-Power Sensor Processing<\/b><br><script>var project769=\"Many embedded sensing applications periodically acquire batches of data that are transferred to memory before any processing is performed. This data-transfer time can instead be exploited to perform lightweight inference <strong>as the sensor data enters the system<\/strong>, enabling irrelevant data to be discarded and the main processing system to remain inactive.<p>At the Embedded Systems Laboratory (ESL) of EPFL, we develop X-HEEP, an open-source and highly configurable RISC-V microcontroller platform designed for low-power embedded systems and easy integration of custom hardware accelerators. X-HEEP includes a DMA architecture featuring a FIFO-based interface for tightly coupling streaming accelerators.<\/p><p><strong>The project aims to investigate lightweight streaming wake-up models and exploit the DMA streaming interface to process sensor data directly during acquisition.<\/strong><\/p><p>The work will first explore, through literature study and Python experiments, which models can operate under strict streaming constraints. The student will investigate small 1D-convolutional pipelines composed of convolution, activation, aggregation and classification, as well as classical approaches such as matched filters. Representative applications, such as IMU-based event detection and voice activity detection (VAD), will be used to compare these approaches. Selected algorithms will then be implemented on the X-HEEP RISC-V CPU to establish a baseline. Based on the results, the project can be extended to the design of a DMA-coupled streaming accelerator capable of executing the selected wake-up pipeline while sensor data is transferred into the system.<\/p><p>This research internship will be carried out at the ESL at EPFL. The student will be under the supervision of Tommaso Terzano, Dr. David Mallas&eacute;n Quintana, and Prof. David Atienza.<\/p><p><strong>Throughout the project, the student will learn:<\/strong><\/p><ul> <li>About always-on and wake-up processing for low-power sensing systems.<\/li> <li>How to design and evaluate extremely lightweight machine-learning models.<\/li> <li>How streaming constraints affect neural-network and classical signal-processing algorithms.<\/li> <li>How DMA engines and streaming hardware accelerators can cooperate during sensor acquisition.<\/li> <li>How to approach hardware\/software co-design from algorithmic exploration to accelerator design.<\/li> <\/ul><p><strong>Project objectives:<\/strong><\/p><ul> <li>Study the literature on lightweight wake-up processing, streaming neural networks, and classical event-detection approaches.<\/li> <li>Investigate small 1D-convolutional wake-up models composed of convolution, activation, aggregation, and classification.<\/li> <li>Investigate classical alternatives, particularly matched-filter\/FIR-based detection, and compare their suitability for streaming execution.<\/li> <li>Implement and evaluate representative approaches in Python.<\/li> <li>Study the trade-offs between detection performance, computational complexity, memory requirements, and streamability.<\/li> <li>Evaluate representative applications such as IMU-based event detection and voice activity detection.<\/li> <li>Implement selected algorithms in C and deploy them on the X-HEEP RISC-V CPU to establish a software baseline.<\/li> <li>Characterize their execution cost and identify the operations that can be performed directly while sensor data is being transferred by the DMA.<\/li> <li>Define a streaming architecture capable of processing the complete wake-up pipeline through the DMA FIFO interface.<\/li> <li>[IF TIME ALLOWS] Implement and integrate a custom streaming accelerator with the X-HEEP DMA.<\/li> <\/ul><p><strong>Required knowledge and skills:<\/strong><\/p><ul> <li>Good Python programming skills<\/li> <li>Good C programming skills<\/li> <li>Excellent knowledge of machine learning and digital signal processing<\/li> <li>Basic knowledge of computer architecture and embedded systems<\/li> <li>Confidence working with Linux systems and Git<\/li> <li>Interest in low-power edge AI and hardware\/software co-design<\/li> <li>Familiarity with RISC-V, RTL design, or hardware accelerators is a plus<\/li> <li>Autonomy and scientific curiosity<\/li> <\/ul><p><strong>Type of work:<\/strong><\/p><p>35% Research and algorithm exploration<br \/>30% Python modelling and evaluation<br \/>25% embedded implementation and benchmarking<br \/>10% hardware exploration<\/p><p><strong>Lab:<\/strong> ESL&nbsp; <br \/><strong>Section:<\/strong> SEL&nbsp; <br \/><strong>Supervisors:<\/strong> Tommaso Terzano, Dr. David Mallas&eacute;n Quintana, Prof. David Atienza<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Tommaso Terzano, Dr. David Mallas\u00e9n Quintana, Prof. David Atienza<br> Contact email: <a href='mailto:tommaso.terzano@epfl.ch; david.mallasen@epfl.ch; david.atienza@epfl.ch?subject=Streaming Wake-Up Networks for Low-Power Sensor Processing'>tommaso.terzano@epfl.ch; david.mallasen@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project769minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Tommaso Terzano, Dr. David Mallas\u00e9n Quintana, Prof. David Atienza<br>\";<\/script>\n<span id=project769><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Tommaso Terzano, Dr. David Mallas\u00e9n Quintana, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project769',project769); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><a href=#_ onclick=opendesc('project761',project761); style='position:relative;z-index:99;'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/project761.png><\/a><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor761><\/a><b><span style='font-size: 20px;'>Development of a patch-sized PCB for long-term skin conductance monitoring using HEEPidermis<\/b><br><script>var project761=\"<p><span style='color: #000000'>Galvanic Skin Response (GSR) is a non-invasive modality that allows us to measure emotional triggers -such as stress- by measuring the skin&rsquo;s response to a superficial electrical current. Developing miniaturized hardware that enables ambulatory monitoring of GSR in an inconspicuous way would open the door to major research on stress focused on first responders and victims of violence. However, the integration of a complex system into a wearable device requires careful co-design between hardware and software centered at power\/area efficiency and user comfort.  <\/span> <\/p> <p><span style='color: #000000'>The Embedded Systems Laboratory (ESL) of EPFL has been at the forefront of developing wearable solutions for healthcare monitoring. With vast experience in hardware, firmware and application design, ESL has distinguished itself for tackling challenges in a holistic manner introducing innovations at every layer of the stack. <\/span> <\/p> <p><span style='color: #000000'><strong>The project aims to develop a fully functional patch-sized prototype for GSR recording based on the HEEPidermis chip of ESL.<\/strong><\/span> <\/p> <p><span style='color: #000000'>The work includes challenges on rigid-flex PCB design and integration into a flexible packaging, pruning of required components, verification of signal integrity and development of a wireless transmission stack (from firmware to antenna design, encompassing the design or selection of the receiving hardware). Furthermore, the project can be expanded into integrating wireless battery charging to extend the lifespan of an hermetically-sealed device. <\/span> <\/p> <p><span style='color: #000000'>This <strong>Research Assistantship<\/strong> opportunity will be carried out at the ESL at EPFL. ESL is an active group (22 Ph.D. students among 45 members) involved in many research aspects. The student will be under the supervision of Juan Sapriza and Prof. David Atienza.<\/span> <\/p>  <h3><span style='color: #000000'>Throughout the project, the student will learn:<\/span><\/h3> <ul>     <li>         <p><span style='color: #000000'>About the relevance and complexity of acquisition of biological signals through wearable systems, in particular of GSR. <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>How to design and develop wearable-oriented electronics (datasheets, schematics, low power ICs, etc).<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>To conduct experiments with biological signals, employing rigorous methods to ensure data accuracy and reproducibility.<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>To develop miniaturized, rigid-flex PCBs, respecting strict packaging and usability requirements.<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>To design and test RF circuits made from discrete components. <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>To program custom-made MCU systems in a resource-constrained environment. <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>To work with a multidisciplinary team of people all contributing to the same project.<\/span>         <\/p>     <\/li> <\/ul> <h3><span style='color: #000000'>Project objectives:<\/span><\/h3> <ul>     <li>         <p><span style='color: #000000'>Understand the context and requirements of wearable, ambulatory GSR measurement. <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Test and characterize the first version of a standalone HEEPidermis PCB<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Using the available SDK, acquire data and transmit it through a wired connection.<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Propose and implement changes and improvements to the PCB to ensure a comfortable and reliable recording.<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Propose and iterate over a form-factor for inconspicuous long-term recording. <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Validate and re-design the prototype for wireless data transfer.<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Develop or select the hardware of the data receiver<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Document the designs and experiment procedures. <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>[IF TIME ALLOWS] Design a wireless power transfer circuit to charge the system&rsquo;s battery. <\/span>         <\/p>     <\/li> <\/ul> <h3><span style='color: #000000'>Required knowledge and skills:<\/span><\/h3> <ul>     <li>         <p><span style='color: #000000'>Advanced PCB design<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Knowledge on discrete-hardware design and optimization<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Understanding of RF circuits <\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Intermediate C programming<\/span>         <\/p>     <\/li>     <li>         <p><span style='color: #000000'>Creativity, autonomy and scientific rigor<\/span>         <\/p>     <\/li> <\/ul> <h3><span style='color: #000000'>Type of work: <\/span><\/h3> <p><span style='color: #000000'>10% Research, 30% Design, 30% Implementation, 30% Testing and validation<\/span> <\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Juan Sapriza, Prof. David Atienza<br> \";<\/script>\n<script>var project761minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Juan Sapriza, Prof. David Atienza<br>\";<\/script>\n<span id=project761><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Juan Sapriza, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project761',project761); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor676><\/a><b><span style='font-size: 20px;'>Automation of a semi-custom design and verification flow for ultra-low-leakage always-on circuits relying on differential logic<\/b><br><script>var project676=\"<div class='entry-content mb-5'> \t\t <p>Microcontrollers (MCUs) are used in a wide range of applications ranging from sensor monitoring to robotics and automotive.<\/p> <p>Thanks to their versatility, MCUs are typically chosen as edge computing platforms.&nbsp;<\/p> <p>IoT, wearable, and edge-computing applications are typically profiled in 4 different phases:<\/p> <ol><li>acquisition<\/li><li>pre-processing<\/li><li>processing<\/li><li>transmission or stimulation via actuators<\/li><\/ol> <p>Acquisitions and stimulation can be performed either with external  analog-digital (ADC) and digital-analog (DAC) converters, or with  integrated ones.<\/p> <p>For the latter, the processor (usually implemented in digital) and  the analog components are implemented in the same technology node and  integrated in the same system-on-chip (SoC), requiring  analog-digital-interfaces.<\/p> <p>With a wide collaboration between EPFL, Imperial College London,  Universidad Carlos III de Madrid, and Politecnico di Torino, we are  building HEEPidermis: an SoC that integrates both the processing  elements and the acquisition and stimulation required to obtain  precision measurements of impedance and conductance of the skin in a  low-power and autonomous manner.&nbsp;<\/p> <p>The processor is based on X-HEEP, an open-source RISC-V configurable  and extendable microcontroller, and includes smart pre-processing of  data coming from ADCs, while the acquisition and stimulus components are  based on a set of possible ADCs (VCO-based, &Delta;&Sigma; and Level-Crossing), and  a current DAC, respectively. HEEPidermis also includes an FLL and LDO  to reduce the required count of off-chip components.&nbsp;<\/p> <p>The full chip is going to be implemented into the TSMC 65 LP technology.&nbsp;<\/p> <p>This project proposes to:<\/p> <ul><li>Do the full-custom layout of Analog\/Digital\/Mixed-Signal blocks.  These include ADCs and DACs, as well as mixed-signal components. <ul><li>Specifications and schematics will be provided<\/li><li>The layout will be done by using Cadence Virtuoso <ul><li>Possibly requiring modifying the Schematic<\/li><\/ul> <\/li><li>The layout must be equivalent to the schematic (LVS) and performed with Calibre<\/li><li>The layout must be DRC-free and performed with Calibre<\/li><li>Simulations with the parasitic extracted netlist will need to be performed to evaluate the final design<\/li><li>LEF and LIB need to be generated so that such IPs can be integrated into the digital-on-top flow used to build HEEPidermis<\/li><\/ul> <\/li><li>Design level shifters that interface the asynchronous-analog side with the synchronous-digital side.&nbsp; <ul><li>Specifications will be provided &ndash; as well as a baseline schematic and layout<\/li><li>The schematic and the layout will be done by using Cadence Virtuoso<\/li><li>The layout must be equivalent to the schematic (LVS) and performed with Calibre<\/li><li>The layout must be DRC-free and performed with Calibre<\/li><li>Simulations with the parasitic extracted netlist will need to be performed to evaluate the final design<\/li><li>LEF and LIB need to be generated so that such IPs can be integrated into the digital-on-top flow used to build HEEPidermis<\/li><\/ul> <\/li><li>Perform the back-end of the digital and analog components of HEEPidermis using Cadence Innovus <ul><li>LEF and LIB of the analog\/mixed-signal components are generated by the tasks above<\/li><li>the processor netlist and constraints are provided<\/li><li>AMS verification of the whole HEEPidermis will be performed together with the rest of the team<\/li><\/ul> <\/li><\/ul> <p>The project will be carried out at the ESL at EPFL, one of the world&rsquo;s top-class universities.<\/p> <p><strong>Required knowledge and skills:<\/strong><\/p> <ul><li>Synopsys Design Compiler<\/li><li>Cadence Innovus and Virtuoso<\/li><li>Siemens Calibre<\/li><li>Abstraction files syntax of LEF, LIB, etc.<\/li><li>Spice\/HSpice and SystemVerilog\/Verilog<\/li><li>Good analytical skills<\/li><li>Teamwork and git<\/li><\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English&nbsp;<\/li><\/ul> <p><br \/><strong>Type of work:<\/strong> 10% theory analysis, 90% design and simulation<\/p> \t<\/div>                         <div class='post-nav py-md-1'>                                 <div class='nav-prev'>         <\/div><\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Alexandre Levisse, Prof. Matias Miguez, Robin Leplae, Juan Sapriza, Prof. David Atienza<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch;alexandre.levisse@epfl.ch;matias.miguez@epfl.ch;robin.leplae@epfl.ch;juan.sapriza@epfl.ch;david.atienza@epfl.ch?subject=Automation of a semi-custom design and verification flow for ultra-low-leakage always-on circuits relying on differential logic'>davide.schiavone@epfl.ch;alexandre.levisse@epfl.ch;matias.miguez@epfl.ch;robin.leplae@epfl.ch;juan.sapriza@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project676minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Alexandre Levisse, Prof. Matias Miguez, Robin Leplae, Juan Sapriza, Prof. David Atienza<br>\";<\/script>\n<span id=project676><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Alexandre Levisse, Prof. Matias Miguez, Robin Leplae, Juan Sapriza, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project676',project676); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><tr><td colspan=2><h3> Master Projects<br><br><\/h3><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor768><\/a><b><span style='font-size: 20px;'>Characterisation and Integration of a Modular Optical Physiological Sensing Module into VersaSens<\/b><br><script>var project768=\"<p>The objective of this project is to characterise and integrate a modular optical physiological sensing module developed for wearable measurement of cardiovascular and muscle-oxygenation parameters into the VersaSens platform.<\/p> <p>VersaSens is a modular, multimodal, extendable and reconfigurable Edge-AI wearable platform developed at the Embedded Systems Laboratory (ESL). Its modular architecture enables the integration of additional sensing technologies and application-specific hardware, providing a flexible platform for exploring new wearable sensing modalities.<\/p> <p>The project will investigate the integration of the FI-SmO\u2082 module, a compact optical sensing extension providing PPG and NIRS measurement capabilities, into VersaSens. The work will be conducted in two main phases.<\/p> <p>In the <strong>first phase<\/strong>, the student will work with the FI-SmO\u2082 module in its native development environment. The objective is to establish a detailed understanding of the sensing architecture, existing software interface and available measurement modes, and to experimentally characterise the system under representative operating conditions. The student will investigate the influence of configurable sensing parameters on measurement quality and characterise the operating range and measurement behaviour of the available sensing modes under representative conditions.<\/p> <p>This stage will establish a validated and documented driver API that can be adapted for integration with VersaSens.<\/p> <p>In the <strong>second phase<\/strong>, the FI-SmO\u2082 module and its validated basic driver interface will be integrated into the VersaSens platform through an appropriate hardware and firmware interface. The student will implement the required VersaSens-side integration, enabling the platform to configure the module, acquire sensor data and expose the measurements to higher-level applications.<\/p> <p>The complete integration will be experimentally validated using VersaSens hardware. A demonstration application will be developed to visualise and monitor the acquired physiological data and verify reliable operation of the integrated sensing module.<\/p> <p>The project therefore combines <strong>optical sensing, embedded systems, device-driver development, hardware integration and experimental characterisation<\/strong>, providing a practical study of how a specialised wearable sensing module can be characterised and integrated into a modular research platform.<\/p> <p><strong>Mandatory tasks<\/strong><\/p> <p>Completion of all mandatory tasks is required to pass the project and obtain a grade of 4. Failure to complete any mandatory task will result in <strong>no pass.<\/strong><\/p> <p><strong>Phase 1 &mdash; FI-SmO\u2082 characterisation and software interface<\/strong><\/p> <ol> <li>Become familiar with the VersaSens platform and the FI-SmO\u2082 sensing module, including their hardware architecture, communication interfaces and practical operation.<\/li> <li>Become familiar with the principles of optical PPG and NIRS-based physiological measurements relevant to the sensing module.<\/li> <li>Study and understand the existing embedded software and sensor interfaces provided for the FI-SmO\u2082 module.<\/li> <li>Experimentally characterise the sensing module under representative measurement conditions and investigate the influence of relevant configurable parameters on signal quality, dynamic range, stability and measurement reliability.<\/li> <li>Evaluate the available measurement modes and characterise their operating ranges and measurement behaviour under representative conditions.<\/li> <li>Review, clean and structure the existing device-driver implementation into a modular driver with clearly defined functions for hardware initialisation, sensor hub configuration\/s, measurement-mode control and data acquisition.<\/li> <li>Develop and document a basic, well-defined driver API for configuring the FI-SmO\u2082 module and acquiring the relevant sensor measurements, suitable for subsequent adaptation to the VersaSens firmware environment.<\/li> <\/ol> <p><strong>Phase 2 &mdash; VersaSens integration<\/strong><\/p> <ol> <li>Design and implement a simple interface\/adapter PCB to connect the FI-SmO\u2082 module to the VersaSens platform.<\/li> <li>Adapt the validated FI-SmO\u2082 module driver interface for the VersaSens firmware environment and implement the required VersaSens-side integration.<\/li> <li>Implement reliable communication and acquisition of the relevant sensor data through VersaSens.<\/li> <li>Test and validate the complete FI-SmO\u2082&ndash;VersaSens integration using the prototyped hardware.<\/li> <li>Develop a demonstration application capable of visualising, monitoring and validating real-time acquisition from the integrated sensing module.<\/li> <li>Prepare a well-documented development package uploaded on lab GIT repository including firmware source code, software libraries, configuration information, host-side tools, readme.md, and technical documentation.<\/li> <\/ol> <p><strong>Optional tasks<\/strong><\/p> <p>Once all mandatory tasks have been completed, the project may be extended with one or more of the following tasks. Each completed optional task contributes<strong> 0.5 points <\/strong>to the final grade, up to a<strong> maximum of 2 additional points.<\/strong><\/p> <ol> <li>Develop a systematic calibration procedure for evaluating and adapting optical sensing configurations under different measurement conditions.<\/li> <li>Develop an automated procedure for evaluating and selecting suitable sensing configurations based on predefined measurement-quality criteria.<\/li> <li>Characterise the power&ndash;performance trade-off of alternative sensing configurations and identify configurations offering an appropriate balance between measurement quality and power consumption.<\/li> <li>Extend the VersaSens integration to expose additional sensing functionality or configurable parameters of the FI-SmO\u2082 module through the VersaSens firmware.<\/li> <li>Develop automated characterisation and analysis tools for collecting, comparing and visualising measurements obtained under different sensing configuration<\/li> <\/ol> <p>&nbsp;<\/p> <p><strong>Type of work<\/strong><\/p> <ul> <li><strong>10%<\/strong> VersaSens and FI-SmO\u2082 familiarisation; optical physiological measurement principles<\/li> <li><strong>30%<\/strong> Sensing-system characterisation and operating-parameter evaluation<\/li> <li><strong>30%<\/strong> Device-driver development, interface PCB and VersaSens integration<\/li> <li><strong>20%<\/strong> Testing, validation and demonstration<\/li> <li><strong>10%<\/strong> Documentation and development deliverables<\/li> <\/ul> <p>&nbsp;<\/p> <p><strong>Desired skills<\/strong><\/p> <ul> <li>Strong C\/C++ programming skills for embedded systems (Arduino; RTOS experience is a plus)<\/li> <li>Understanding of communication protocols such as I&sup2;C, SPI and serial interfaces<\/li> <li>Basic PCB design and hardware prototyping skills<\/li> <li>Basic soldering and electronics debugging skills<\/li> <li>Data acquisition and signal processing (Python and\/or MATLAB)<\/li> <li>Interest in optical sensing and physiological measurements<\/li> <li>Experience with embedded systems and RTOS such as FreeRTOS\/Zephyr is advantageous<\/li> <li>Experience with Git version control<\/li> <li>Strong debugging and problem-solving skills<\/li> <\/ul> <p><strong>Soft skills<\/strong><\/p> <ul> <li>Scientific curiosity and rigor<\/li> <li>Strong analytical and experimental mindset<\/li> <li>Ability to work independently<\/li> <li>Good communication skills<\/li> <li>Attention to detail<\/li> <li>Advanced English<\/li> <\/ul> <p><strong>Reference<\/strong><\/p> <p>[1] T. Aminosharieh Najafi et al., &ldquo;VersaSens: An Extendable Multimodal Platform for Next-Generation Edge-AI Wearables,&rdquo; <em>IEEE Transactions on Circuits and Systems for Artificial Intelligence<\/em>, 2024.<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Harshal Arun Sonar, Dr. Taraneh Aminosharieh Najafi, Prof. David Atienza<br> Contact email: <a href='mailto:harshal.sonar@epfl.ch;taraneh.aminoshariehnajafi@epfl.ch;david.atienza@epfl.ch?subject=Characterisation and Integration of a Modular Optical Physiological Sensing Module into VersaSens'>harshal.sonar@epfl.ch;taraneh.aminoshariehnajafi@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project768minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Harshal Arun Sonar, Dr. Taraneh Aminosharieh Najafi, Prof. David Atienza<br>\";<\/script>\n<span id=project768><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Harshal Arun Sonar, Dr. Taraneh Aminosharieh Najafi, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project768',project768); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><a href=#_ onclick=opendesc('project766',project766); style='position:relative;z-index:99;'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/project766.png><\/a><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor766><\/a><b><span style='font-size: 20px;'>Design of a Multiple-Issue Execution Pipeline for an Out-Of-Order RISC-V Processor<\/b><br><script>var project766=\"<h2><span style='color: #000000; font-size: 18px'>Project Description<\/span><\/h2> <p><span style='color: #000000'>Modern edge-computing workloads increasingly rely on heterogeneous platforms where domain-specific tightly-coupled coprocessors, such as vector or tensor accelerators, execute computationally intensive tasks. However, offloading long-latency operations often causes conventional in-order host CPUs to stall, leading to severe resource underutilization and disruption of the data-independent control flow.<\/span> <\/p> <p><span style='color: #000000'>The LEN5 processor, developed at the VLSI Laboratory at Politecnico di Torino and the Embedded Systems Laboratory (ESL) at EPFL, is a modular, dynamically scheduled out-of-order 64-bit RISC-V CPU core. While LEN5 successfully implements dynamic scheduling and dynamic instruction dispatch based on Tomasulo's algorithm, its core microarchitecture currently operates as a functional yet single-issue pipeline.<\/span> <\/p> <p><span style='color: #000000'><strong>This project aims to extend the LEN5 out-of-order RISC-V core into a modular, parameterized superscalar architecture, redesigning the pipeline stages to enable multi-issue fetch and dispatch, extending the execution stage with multi-channel Common Data Buses and parallel execution units, and implementing an advanced basic-block-based out-of-order commit unit with efficient speculative recovery mechanisms.<\/strong><\/span> <\/p> <p><span style='color: #000000'>ESL has extensive experience in embedded systems and microarchitectural research, offering access to cutting-edge tools and resources. The laboratory follows a holistic approach, introducing innovations across multiple layers of the system stack.<\/span> <\/p> <p><span style='color: #000000'>The work will include SystemVerilog RTL design, unit-level and full-core verification, and the use of generative AI tools to accelerate testbench development and coverage analysis. This will enable the student to explore the interaction between microarchitectural design choices and superscalar performance in an out-of-order RISC-V core.<\/span> <\/p><p><span style='color: #000000'>This Master Thesis Project will be carried out at ESL at EPFL. ESL is an active research group working across a broad range of embedded systems and microarchitectural research topics. The student will be supervised by Dr. Michele Caon and Prof. David Atienza.<\/span> <\/p><div><blockquote>     <p><span style='color: #cc0000; font-size: 13px'><strong><em>See the attached PDF for project objectives and evaluation system.<\/em><\/strong><\/span>     <\/p> <\/blockquote><\/div><div><strong><em><br \/><\/em><\/strong><\/div><div><strong><em><br \/><\/em><\/strong><\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Dr. Michele Caon, Prof. David Atienza<br> \";<\/script>\n<script>var project766minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Dr. Michele Caon, Prof. David Atienza<br>\";<\/script>\n<span id=project766><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Dr. Michele Caon, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project766',project766); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=https:\/\/drive.google.com\/file\/d\/1cA3I-M4sF5LBiU7zZz3gCtDbH1rvGEkc\/view?usp=drive_link title='document link'><img width=30 border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/doclink.gif hspace=2 alt='document link'><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor764><\/a><b><span style='font-size: 20px;'>Multi-CGRA Tile Accelerator: Integration, Modeling, and Scalability Analysis<\/b><br><script>var project764=\"<p>This project focuses on the implementation and analysis of a <strong>multi-CGRA tile accelerator<\/strong>. You will integrate a DMA engine with multiple CGRA instances and a shared L2 scratchpad, enabling the execution of workloads that are too large or computationally demanding for a single accelerator. The resulting architecture will be evaluated in terms of performance, scalability, and system-level bottlenecks.<\/p> <p>A CGRA consists of a mesh of tightly coupled Processing Elements (PEs) that can execute operations in parallel and exchange results with neighboring elements. By efficiently mapping application kernels onto this mesh, CGRAs provide an attractive balance between energy efficiency, performance, and flexibility. However, the increasing computational and data-movement requirements of applications such as radio astronomy, genomics, and machine learning motivate architectures that can scale beyond a single CGRA.<\/p> <p>To enable <strong>scale-up through accelerator replication<\/strong>, the project will combine multiple independently controlled CGRAs with a DMA and shared scratchpad memory. The complete tile will be implemented in RTL and modeled in C++, with the system-level model validated against RTL simulations. Representative workloads will be deployed across multiple CGRAs to characterize scalability, utilization, memory traffic, and the main performance bottlenecks of the architecture.<\/p> <p><strong>Tasks description<\/strong><\/p> <ol><li><strong>Understand the CGRA and DMA<\/strong><br \/>Study the existing CGRA and DMA implementations, including their interfaces, control mechanisms, memory accesses, and available workloads. Understand how data is transferred between the CGRAs, DMA, L2 scratchpad, and external system. <\/li><li><strong>C++ simulation framework<br \/><\/strong>Develop a system-level C++ model of a tile containing multiple CGRA instances, a DMA, and an L2 scratchpad. Model accelerator execution, data transfers, and shared-resource contention. Validate and calibrate the model using RTL simulation results.<\/li><li><strong>RTL integration and validation<\/strong><br \/>Integrate the DMA, L2 scratchpad, and multiple CGRA instances into a complete RTL tile. Develop the required interconnect and control logic, validate the implementation with representative workloads, and compare its behavior against the C++ model.<\/li> <li><strong>Performance and scalability analysis<\/strong><br \/>Deploy the provided workloads across multiple CGRA instances and evaluate execution time, speedup, utilization, and data-movement overhead. Identify the main limits to scalability, including DMA throughput, memory bandwidth, synchronization, and shared-resource contention.<\/li><\/ol> <p><strong>Project objectives<\/strong><\/p> <p>The fulfillment of the mandatory objectives will result in a passing grade of 4.0. The completion of the two optional objectives will increase the grade by 1.0 point each.<\/p> <ul><li>[Mandatory] Integrate the DMA, L2 scratchpad, and multiple CGRA instances and validate the RTL implementation.<\/li><li>[Mandatory] Develop a C++ model of the tile and validate and characterize it against RTL simulation.<\/li><li>[Mandatory] Deploy <u>at least one<\/u> of the provided workloads across multiple CGRA instances and characterize its performance. The deployment of a workload on a single CGRA instance will be provided and it&rsquo;s already automated.<\/li><li>[Mandatory] Analyze scalability and identify the main system-level performance bottlenecks and trade-offs.<\/li><li>[Optional] Deploy <u>all three<\/u> provided workloads (FFT, GEMM, and convolution) across multiple CGRAs and characterize their performance.<\/li><li>[Optional] Deploy the CGRA tile on an FPGA and use Linux running in the hard-ip cores of the device to transfer data from external DDR.<\/li><\/ul> <p><strong>Required knowledge and skills<\/strong><\/p> <ul><li>Proficiency in RTL design and programming (VHDL, Verilog, or SystemVerilog).<\/li><li>Advanced knowledge of C\/C++ programming and hardware modeling.<\/li><li>Basic understanding of computer architecture and hardware\/software interfaces.<\/li><li>Familiarity with DMA, OBI, and AXI-based systems is beneficial.<\/li><li>Experience with FPGA design and implementation is beneficial.<\/li><li>Strong analytical thinking and scientific curiosity.<\/li><\/ul> <p><strong>References<\/strong><\/p> <p>[1] Rodr&iacute;guez &Aacute;lvarez, Rub&eacute;n, et al. &quot;An open-hardware coarse-grained reconfigurable array for edge computing.&quot; <em>Proceedings of the 20th ACM International Conference on Computing Frontiers<\/em>. 2023.<\/p> <p>[2] Benz, Thomas, et al. &quot;A high-performance, energy-efficient modular DMA engine architecture.&quot; <em>IEEE Transactions on Computers<\/em> 73.1 (2023): 263-277.<\/p> <p><strong>Type of work<\/strong><\/p> <ul><li>40% HW design: Integration of the DMA, CGRAs, L2 memory, and interconnect.<\/li><li>40% C++ modeling: Development and validation of the system-level multi-CGRA model.<\/li><li>20% Performance evaluation: Analysis of scalability, utilization, memory traffic, and system-level bottlenecks.<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Denisa Constantinescu, Dr. Miguel Pe\u00f3n Quir\u00f3s, Dr. Giovanni Ansaloni, Prof. David Atienza<br> \";<\/script>\n<script>var project764minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Denisa Constantinescu, Dr. Miguel Pe\u00f3n Quir\u00f3s, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project764><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Denisa Constantinescu, Dr. Miguel Pe\u00f3n Quir\u00f3s, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project764',project764); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor763><\/a><b><span style='font-size: 20px;'>Acquisition-Application co-optimization of a Level-Crossing ADC for Implantable Neural Interfaces<\/b><br><script>var project763=\"<h2><span style='color: #000000'>Project Description<\/span><\/h2> <p><span style='color: #000000'>Neural signals are information-rich modalities that can help guide the diagnosis and treatment of a wide range of neurological disorders. The highest-quality recordings are obtained directly from the surface or inner layers of the brain, but the invasiveness of these approaches imposes stringent area and energy constraints on acquisition and processing hardware.<\/span> <\/p> <p><span style='color: #000000'>To reduce the burden associated with the high-bandwidth, high-dynamic-range data acquisition required in these applications, Level-Crossing (LC) Analog-to-Digital Converters (ADCs) are gaining increasing attention. Unlike conventional ADCs, LC-ADCs generate samples only when the input signal crosses predefined amplitude levels, potentially reducing the output data rate by orders of magnitude. However, this event-driven acquisition strategy departs from traditional uniformly sampled systems, and established methods for configuring LC-ADCs and adapting downstream processing algorithms are still lacking.<\/span> <\/p> <p><span style='color: #000000'><strong>This project aims to build upon an LC-ADC already developed at the Embedded Systems Laboratory (ESL) of EPFL and iteratively co-optimize the acquisition hardware and downstream applications to achieve the best overall trade-off between power, area, and application performance.<\/strong><\/span> <\/p> <p><span style='color: #000000'>ESL has extensive experience in developing technologies for healthcare monitoring, spanning hardware, firmware, and application design. The laboratory follows a holistic approach, introducing innovations across multiple layers of the system stack.<\/span> <\/p> <p><span style='color: #000000'>The work will include transistor-level simulations, circuit modifications, and potentially circuit redesign. Neural signals will be processed using different ADC configurations, with hardware optimization guided by the performance obtained on relevant medical tasks and applications. This will enable the student to explore the interaction between circuit-level design choices, signal representation, and application-level performance.<\/span> <\/p><p><span style='color: #000000'>This Master Thesis Project will be carried out at ESL at EPFL. ESL is an active research group with 45 members, including 22 Ph.D. students, working across a broad range of embedded systems research topics. The student will be supervised by Juan Sapriza and Prof. David Atienza.<\/span> <\/p><div><blockquote>     <p><span style='color: #cc0000'><strong><em>See the attached PDF for project objectives and evaluation system.<\/em><\/strong><\/span>     <\/p> <\/blockquote><\/div><div><strong><em><br \/><\/em><\/strong><\/div><div><strong><em><br \/><\/em><\/strong><\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Juan Sapriza, Prof. David Atienza<br> \";<\/script>\n<script>var project763minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Juan Sapriza, Prof. David Atienza<br>\";<\/script>\n<span id=project763><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Juan Sapriza, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project763',project763); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><a href=https:\/\/docs.google.com\/document\/d\/1cfvgrHzUvGqlnA4ZF2AWcOXOQSD05iYtfc4qgGdfcPM\/edit?usp=sharing title='document link'><img width=30 border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/doclink.gif hspace=2 alt='document link'><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor745><\/a><b><span style='font-size: 20px;'>Precision-Scalable Systolic Arrays for High-Efficiency LLM Inference Acceleration<\/b><br><script>var project745=\"<p>This project focuses on the design and implementation of a scalable-bitwidth systolic array, a specialized hardware accelerator aimed at improving the computational efficiency of modern artificial intelligence workloads. Its primary objective is to develop an architecture that can dynamically support varying numerical precisions through a template-based design, ranging from low-bit quantization (e.g., INT2\/INT3\/INT4) to higher-precision formats such as FP16.<\/p>    <p>A systolic array consists of a network of tightly coupled data-processing units, referred to as Processing Elements (PEs), through which data flows in a structured manner. Each PE performs a small portion of the overall computation&mdash;typically multiply-accumulate (MAC) operations&mdash;and forwards intermediate results to neighboring elements. Systolic arrays architecture enables massive parallelism and high data reuse. By significantly reducing accesses to off-chip memory, they achieve very high energy efficiency for matrix multiplication workloads, which form the computational backbone of deep learning and large language models.<\/p>    <p>Despite their efficiency, state-of-the-art systolic arrays are typically limited to a single, fixed bitwidth. Most existing designs target a specific numerical format, such as INT8 or FP16, which makes them less adaptable to the diverse precision requirements of modern workloads. As a result, these architectures struggle to efficiently support workloads with varying numerical characteristics, limiting their flexibility and broader applicability.<\/p>    <p>To address this limitation, this project proposes a reconfigurable, scalable-bitwidth systolic array. To this end, the student will be tasked with the development of a template-based PE architecture that enables the generation of practical systolic arrays supporting varying bitwidths. Multiple PEs can be dynamically grouped to form scalable matrix-multiplication accelerators with configurable precision. This approach aims to bridge the gap between flexibility and efficiency, enabling hardware accelerators that better match the diverse precision demands of modern AI models.<\/p>    <p><strong>Tasks description<\/strong><\/p>    <ol class='wp-block-list'> <li>Understand the architecture of the TiC-SAT systolic array under development at ESL_EPFL, including its PE design and interconnects.<\/li>    <li>Implement a PE design template that supports various bitwidth configuration.<\/li>    <li>Generate and test a heterogeneous systolic array hardware designs using the scalable-bitwidth PEs to enable flexible precision support.<\/li>    <li>Automate the integration of systolic array instances in RISC-V systems-on-chip within the X-HEEP open-hardware framework.<\/li> <\/ol>    <p><strong>Project objectives<br \/><\/strong>The fulfillment of the following objective is required for a passing grade (4.0)<\/p>    <ul class='wp-block-list'> <li>Extend the TiC-SAT PE design to support multiple bitwidth, including a testbench and testsuite to test and validate the design.<\/li>    <li>Create a template-based generator of the systolic array. The template must enable the generation of instances of the systolic array from configuration parameters, specifying the array size, supported bitwidths and the arrangement of PEs.<\/li>    <li>Characterize the runtime latency and energy efficiency of the modified systolic array for at least four different bitwidth configurations.<\/li>    <li>Integrate an instance systolic array design into the XHEEP [2], and test the performance of at least on AI model on the resulting system,<\/li> <\/ul>    <p>The completion of each of the following tasks will add 0.5 extra points to the project grade<\/p>    <ul class='wp-block-list'> <li>Build a functional simulator for the SA in python.<\/li>    <li>Implement the system on an FPGA development board.<\/li>    <li>Automate the template systolic array integration into XHEEP.<\/li>    <li>Explore the rousource\/accuracy trade-offs realized by different systolic array configurations.<\/li> <\/ul>    <p><strong>Required knowledge and skills<\/strong><\/p>    <ul class='wp-block-list'> <li>Proficiency in RTL design and programming (e.g., VHDL or Verilog).<\/li>    <li>Basic understanding of computer architecture.<\/li>    <li>Strong analytical thinking and scientific curiosity.<\/li> <\/ul>    <p><strong>References<\/strong><\/p>    <p>[1] A. Amirshahi, J. Klein, G. Ansaloni, D. Atienza, &ldquo;TiC-SAT: Tightly-coupled Systolic Accelerator for Transformers&rdquo;,&nbsp; 28th Asia and South Pacific Design Automation Conference (ASP-DAC &rsquo;23), Tokyo, Japan, doi: 10.1145\/3566097.3567867<\/p>    <p>[2] S. Machetti, P. D. Schiavone, G. Ansaloni, M. Pe&oacute;n-Quir&oacute;s and D. Atienza, &ldquo;X-HEEP: An Open-Source, Configurable and Extendible RISC-V Platform for TinyAI Applications,&rdquo;&nbsp;<em>2025 IEEE Computer Society Annual Symposium on VLSI (ISVLSI)<\/em>, Kalamata, Greece, 2025, pp. 1-6, doi: 10.1109\/ISVLSI65124.2025.11130281.<\/p>    <p><strong>Type of work<\/strong><\/p>    <ul class='wp-block-list'> <li>80% <strong>HW design<\/strong>: development and implementation of the scalable-bitwidth systolic array hardware.<\/li>    <li>20% <strong>Performance Evaluation<\/strong>: Benchmarking and analysis of various machine learning workloads on the designed hardware.<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> Contact email: <a href='mailto:ruben.rodriguezalvarez@epfl.ch; yuxuan.wang@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch?subject=Precision-Scalable Systolic Arrays for High-Efficiency LLM Inference Acceleration'>ruben.rodriguezalvarez@epfl.ch; yuxuan.wang@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project745minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project745><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project745',project745); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor736><\/a><b><span style='font-size: 20px;'>From Telemetry to Flexibility Services: Schema Design for Data Centres as Energy Prosumers<\/b><br><script>var project736=\"<p>Data centres hosting HPC, cloud, and AI workloads are becoming megawatt-scale actors in electricity systems, yet they are still treated as passive consumers by energy markets. Recent work shows that data centres can act as grid prosumers by shifting or modulating computation to follow renewable availability and grid needs. In practice, however, this potential remains largely untapped because power and energy information is fragmented across servers, clusters, and facilities, with no common abstraction that allows operators to aggregate behaviour or expose flexibility in a verifiable and service-oriented way.<\/p> <p>Today, low-level telemetry such as CPU and GPU power counters, cluster-level utilisation metrics, and facility indicators like total power draw or PUE are collected using heterogeneous tools and formats. For privacy, security, and commercial reasons, these data cannot usually be shared at fine granularity or directly associated with individual jobs or users managed by schedulers such as SLURM or Kubernetes. As a result, data centres lack mechanisms to aggregate power and energy information in a way that preserves confidentiality while still conveying meaningful system-level flexibility to external energy stakeholders.<\/p> <p>The motivation of this thesis is to address the foundational abstraction gap by designing a common, machine-readable schema for power and energy data that spans from servers to the data center level, supporting privacy-preserving aggregation. Compatibility with external efforts and common tools of practice such as SLURM job metadata or Kubernetes pod annotations and reports is treated as a hard requirement, while ensuring that sensitive information is abstracted, anonymised, or aggregated before exposure. By structuring historical energy data in this way, the schema enables the derivation of auditable flexibility descriptors that can inform flexible Service-Level Agreements (f-SLAs), laying the groundwork for data centers to operate as trustworthy grid prosumers.<\/p> <p>&nbsp;<\/p> <h1>PROJECT OBJECTIVES<\/h1> <p>OBLIGATORY OBJECTIVES&nbsp;<\/p> <p>O1. <strong>Define a Common Multi-Scale Power and Energy Data Schema<\/strong>:<br \/>Design a unified, machine-readable schema that captures power and energy information at the server, cluster, and data-centre levels. The schema must define clear semantics, units, temporal resolution, aggregation rules, and provenance metadata, and must be suitable for representing historical energy behaviour in heterogeneous HPC and AI infrastructures.<\/p> <p>O2. <strong>Ensure Mandatory Compatibility with SLURM or Kubernetes<\/strong>:<br \/>Demonstrate that the schema can be instantiated and used in at least one production-relevant scheduling environment. The schema must be directly mappable to either SLURM job metadata and accounting reports, or Kubernetes job or pod annotations and execution reports, using supported interfaces without modifying core scheduler behaviour.<\/p> <p>O3. <strong>Enable Privacy-Preserving Aggregation for Flexibility Extraction:<\/strong><br \/>Design aggregation mechanisms that transform fine-grained power and energy data into non-sensitive, system-level indicators. The aggregation must prevent disclosure of job- or user-level information while preserving the ability to infer meaningful flexibility characteristics at the cluster and data-centre level.<\/p> <p>STRETCH OBJECTIVES&nbsp;<\/p> <p>S1. <strong>Derive Flexibility Descriptors for Flexible Service-Level Agreements:<\/strong> Translate aggregated historical power and energy data into flexibility descriptors such as load-shifting capacity, power modulation ranges, and temporal elasticity, and show how these descriptors can inform flexible Service-Level Agreements between data centres and energy providers.<\/p> <p>S2. <strong>Cross-Platform Validation Across HPC and AI Environments<\/strong>:&nbsp; Validate the schema and aggregation mechanisms in both a traditional HPC context and a modern AI platform, demonstrating that equivalent information and flexibility descriptors can be derived from SLURM-based and Kubernetes-based environments.<\/p> <p>S3. <strong>Represent, Validate and Present this work within communities of interest<\/strong>: Collaboration with, presentation to and adoption by data center practitioners and researchers from the Energy Efficiency HPC Working Gourp, active scientific data experimental sites and\/or industry partners would be an exceptional metric of success.<\/p> <h1>REQUIRED KNOWLEDGE AND SKILLS<\/h1> <ul><li>Programming skills (Python and\/or C\/C++)<\/li><li>Familiarity with Linux-based systems<\/li><li>\u200b\u200bSoftware engineering practices: version control, modular code design, documentation, and reproducibility<\/li><li>Scientific curiosity and good analytical skills<\/li><li>Basic understanding of batch scheduling or container orchestration<\/li><li>Prior experience with SLURM or Kubernetes is beneficial but not required<\/li><li>Interest in distributed systems, HPC, or AI infrastructure<\/li><li>Interest in energy systems, sustainability, or infrastructure policy<\/li><\/ul> <h1>Type of Work<\/h1> <ul><li><strong>Theoretical Analysis (30%)<\/strong> <ul><li>Formalization of user intent and sustainability signals.<\/li><li>Abstraction of scheduler-independent semantics.<\/li><\/ul> <\/li><li><strong>Design and Implementation (45%)<\/strong> <ul><li>Design of data acquisition schema.<\/li><li>Implementation for SLURM and Kubernetes environments.<\/li><\/ul> <\/li><li><strong>Testing and Evaluation (25%)<\/strong> <ul><li>Validation of correctness, expressiveness, and overhead.<\/li><li>Cross-platform comparison of schema instantiations.<\/li><\/ul> <\/li><\/ul> <p><strong>Expected Outcomes<\/strong><\/p> <ul><li>To pass, the student will provide a well-defined and documented common schema for power and energy data across data-centre scales and a prototype demonstrating how the schema supports flexibility extraction (obligatory objectives)<\/li><li>For the maximum grade, the student will analyse and demonstrate how this schema can contribute to research on data centres as grid prosumers (any two of the stretch objectives)<\/li><li>A Master&rsquo;s thesis report, presentation, and GitHub repository<\/li><\/ul> <h1>SUPERVISION<\/h1> <p>This project will be conducted in a research environment focused on <strong>sustainable computing and large-scale systems<\/strong>, under the supervision of experts in HPC, AI infrastructure, and energy-aware systems:<\/p> <ul><li>Embedded Systems Laboratory (ESL), EPFL: <strong>Dr. Denisa-Andreea Constantinescu<\/strong> &ndash; <a href='mailto:denisa.constantinescu@epfl.ch'>denisa.constantinescu@epfl.ch<\/a>, <strong>Prof. David Attienza<\/strong><\/li><li>Los Alamos National Laboratory, HPC Workload Management: <strong>Steven Senator<\/strong><\/li><\/ul> <h1>References<\/h1> <ul><li>SLURM Workload Manager:<a href='https:\/\/slurm.schedmd.com\/'> https:\/\/slurm.schedmd.com\/<\/a><\/li><li>Kubernetes Documentation:<a href='https:\/\/kubernetes.io\/'> https:\/\/kubernetes.io\/<\/a><\/li><li>EEHPCWG Operational Data Analytics, <a href='https:\/\/hpc-oda-org.pages.dev\/'>https:\/\/hpc-oda-org.pages.dev\/<\/a>&nbsp;<\/li><li>Horace, Leslie A., Christopher Stokes, Craig S. Walker, Anvitha Ramachandran, William M. Jones, Nathan A. DeBardeleben, and Steven T. Senator. &quot;Energy Forecasting in High Performance Computing Datacenters Using Machine Learning.&quot; In 2025 IEEE International Conference on AI and Data Analytics (ICAD), pp. 1-10. IEEE, 2025.<\/li><\/ul> <p>Colangelo, Philip, Ayse K. Coskun, Jack Megrue, Ciaran Roberts, Shayan Sengupta, Varun Sivaram, Ethan Tiao et al. &quot;Turning AI Data Centers into Grid-Interactive Assets: Results from a Field Demonstration in Phoenix, Arizona.&quot; <a href='https:\/\/arxiv.org\/pdf\/2507.00909v1'><em>https:\/\/arxiv.org\/pdf\/2507.00909v1<\/em><\/a> &nbsp;(2025).<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Denisa Constantinescu, Prof. David Atienza, Steven Senator of Los Alamos National Laboratory, HPC Workload Management<br> Contact email: <a href='mailto:denisa.constantinescu@epfl.ch; david.atienza@epfl.ch; sts@lanl.gov?subject=From Telemetry to Flexibility Services: Schema Design for Data Centres as Energy Prosumers'>denisa.constantinescu@epfl.ch; david.atienza@epfl.ch; sts@lanl.gov<\/a><br>\";<\/script>\n<script>var project736minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Denisa Constantinescu, Prof. David Atienza, Steven Senator of Los Alamos National Laboratory, HPC Workload Management<br>\";<\/script>\n<span id=project736><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Denisa Constantinescu, Prof. David Atienza, Steven Senator of Los Alamos National Laboratory, HPC Workload Management<br> <a href=#_ onclick=opendesc('project736',project736); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/urbantwin.ch target=_blank title='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/197.png width=70 alt='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor709><\/a><b><span style='font-size: 20px;'>Eyes on the Unmeasured: Blind Forecasting through Multivariate Attention Networks<\/b><br><script>var project709=\"<p>Large-scale wastewater systems contain many nodes where installing sensors is impractical or too costly. Blind forecasting aims to reconstruct or predict the time evolution of <em>unobserved<\/em> endogenous variables using only the available sensor signals and environmental drivers, such as precipitation. This capability is central to building scalable digital twins [2] that require sparse instrumentation while maintaining high-fidelity, network-wide predictive capabilities.<\/p> <p>Transformers are the state-of-the-art for hydrological forecasting due to their ability to model long-range temporal dependencies and cross-variable interactions. Building on the AquaCast multi-input Transformer [1], this project investigates how such architectures can be extended to perform <strong>blind forecasting<\/strong>&mdash;predicting a target time-series that is not provided as input, but must be inferred from the multivariate structure of the drainage network.<\/p> <p>The key research questions include: Which nodes can be reliably inferred from others? How does hydrodynamic coupling influence blind prediction? How much do exogenous variables such as rainfall contribute to successful reconstruction? The goal is to develop a robust, interpretable blind forecasting module and identify architectural or environmental factors that determine blind predictability.<\/p> <p><strong>TASKS<\/strong><\/p> <ol> <li><strong> Literature Review<\/strong><ul> <li>Perform a literature review on time-series forecasting and blind forecasting, focusing on methods for cross-variable prediction and latent reconstruction.<\/li> <\/ul><\/li>    <li><strong> Dataset Preparation &amp; Statistical Analysis<\/strong><ul> <li>Preprocess and construct the dataset from raw Lausanne wastewater records for blind forecasting scenarios.<\/li> <li>Conduct statistical analysis of time-series (cross-correlation, mutual information, hydrodynamic lags) to assess blind predictability.<\/li> <\/ul><\/li>    <li><strong> Blind Forecasting Model Development<\/strong><ul> <li>Design, implement, and test blind forecasting <strong>multivariate networks<\/strong>, preferably Transformer-based.<\/li> <li>Provide quantitative performance metrics and interpretability of current methods and your design.<\/li> <\/ul><\/li>    <li><strong> Exogenous + Endogenous Blind Forecasting<\/strong><ul> <li>Extend the model to handle blind forecasting with both exogenous and endogenous time-series.<br \/>Conduct a detailed quantitative evaluation and interpretability study.<\/li> <\/ul><\/li>  <\/ol> <p><strong>Optional:<\/strong><\/p> <ul> <li><strong>Robustness of missing samples:<\/strong> Propose embedding approaches to handle missing time-steps even under blind forecasting conditions.<\/li> <li><strong>Missing samples representation:<\/strong> Explore self-supervised reconstruction and evaluate transfer to forecasting tasks.<\/li> <\/ul> <p>&nbsp;<\/p> <p><strong>REQUIREMENTS<\/strong><\/p> <ul> <li>Good Python programming skills.<\/li> <li>Scientific curiosity<\/li> <li>Background in machine learning or signal processing.<\/li> <li>Interest in hydrological systems, digital twins, or representation learning.<\/li> <\/ul> <h3><strong>TYPE OF WORK<\/strong><\/h3> <ul> <li><strong>35% theory<\/strong> (state-of-the-art study, analytical feasibility, cross-sensor dependency analysis).<\/li> <li><strong>65% implementation<\/strong> (architectural design, training pipelines, model evaluation, robustness studies).<\/li> <\/ul> <p><strong>REFERENCES<\/strong><\/p> <p>[1] Abdollahinejad, Golnoosh, Saleh Baghersalimi, Denisa-Andreea Constantinescu, Sergey Shevchik, and David Atienza. &quot;<strong>AquaCast<\/strong>: Urban Water Dynamics Forecasting with Precipitation-Informed Multi-Input Transformer.&quot; <a href='https:\/\/arxiv.org\/abs\/2509.09458'><em>https:\/\/arxiv.org\/abs\/2509.09458<\/em><\/a><\/p> <p>[2] UrbanTwin project: <a href='https:\/\/urbantwin.ch\/'>https:\/\/urbantwin.ch\/<\/a><\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br> Contact email: <a href='mailto:golnoosh.abdollahinejad@epfl.ch; denisa.constantinescu@epfl.ch; david.atienza@epfl.ch?subject=Eyes on the Unmeasured: Blind Forecasting through Multivariate Attention Networks'>golnoosh.abdollahinejad@epfl.ch; denisa.constantinescu@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project709minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br>\";<\/script>\n<span id=project709><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br> <a href=#_ onclick=opendesc('project709',project709); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/urbantwin.ch target=_blank title='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/197.png width=70 alt='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor708><\/a><b><span style='font-size: 20px;'>When Sensors Go Silent: Masked Multivariate Transformers for Robust Urban Water Forecasting<\/b><br><script>var project708=\"<p>Accurate short- and long-horizon forecasting of water dynamics in  urban drainage systems is essential for flood prevention, operational  planning, and digital twin deployments. However, the sensing  infrastructure deployed in real-world wastewater networks often suffers  from unreliable conditions, including battery limitations, harsh  underground environments, radio communication losses, or temporary  sensor outages. As a result, forecasting models must remain functional  when one or more sensors suddenly stop reporting data.<\/p> <p>Recent progress in Transformer-based time-series forecasting&mdash;such as  the AquaCast multi-input architecture developed in our group [1, 2]&mdash;has  demonstrated that attention-based models can effectively combine  endogenous water dynamics signals with exogenous inputs, such as  precipitation history and forecast conditions. Yet, these models  generally assume fully available and clean multivariate input streams.  This assumption breaks down in real settings, where missing channels  compromise the learned temporal and cross-sensor dependencies.<\/p> <p>This Master thesis addresses this gap by developing a <strong>masked Transformer forecasting framework<\/strong>  for urban hydrology applications. The central goal is to enable  accurate and stable forecasting even when any subset of input series is  missing, while preserving the benefits of multi-channel information  fusion. This includes extending AquaCast with masked attention,  missing-aware embeddings, and exogenous&ndash;endogenous fusion strategies.  The project will also explore the limits of robustness and  interpretability under systematic sensor failures.<\/p> <p><strong>TASKS<\/strong><\/p> <p>Toward this end, the student will be responsible for:<\/p> <ol><li><strong> Literature Review<\/strong><ul><li>Conduct a thorough review on time-series forecasting and masked  Transformers, including architectures specifically designed for  channel-dropout, variable masking, or partially observed multivariate  sequences.<\/li><\/ul><\/li>  <li><strong> Dataset Preparation &amp; Statistical Analysis<\/strong><ul><li>Preprocess and construct the dataset from raw Lausanne wastewater records for missing-sensor scenarios.<\/li><li>Perform statistical analysis of the time-series (distributional  characterization, correlations, lag analysis, stationarity, periodicity,  etc.).<\/li><li>Generate controlled masking patterns to simulate sensor failure during training and testing.<\/li><\/ul><\/li>  <li><strong> Masked Multivariate Transformer Architecture<\/strong><ul><li>Design, implement, and test a <strong>masked multivariate Transformer<\/strong> to ensure forecasting remains accurate when one or more sensors stop working.<\/li><li>Provide quantitative evaluation and interpretability analysis  (attention maps, influence of mask tokens, sensitivity to missing  channels).<\/li><\/ul><\/li>  <li><strong> Exogenous + Endogenous Masked Forecasting<\/strong><ul><li>Extend the design to incorporate exogenous and endogenous input time-series together (e.g., rainfall history + forecast).<\/li><li>Evaluate how exogenous data compensates for missing endogenous channels.<\/li><li>Provide quantitative measurements and interpretability as above.<\/li><\/ul><\/li> <\/ol> <p><strong>Optional:<\/strong><\/p> <ul><li><strong>Missing samples robustness:<\/strong> Develop robust embedding strategies to prepare Transformer tokens when time-steps are missing.<\/li><li><strong>Missing samples representation:<\/strong> Investigate self-supervised reconstruction to learn representations transferable to forecasting tasks.<\/li><\/ul> <p><strong>REQUIREMENTS<\/strong><\/p> <ul><li>Strong programming skills in Python (PyTorch preferred).<\/li><li>Background in machine learning, deep learning, or time-series modeling.<\/li><li>Interest in AI robustness, forecasting, and spatio-temporal systems.<\/li><\/ul> <h3><strong>TYPE OF WORK<\/strong><\/h3> <ul><li><strong>35% theory<\/strong> (literature review on masked Transformers, missing-data modeling, experimental design, interpretability).<\/li><li><strong>65% implementation<\/strong> (dataset engineering, architecture development, training, evaluation, and result analysis).<\/li><\/ul> <p><strong>REFERENCES<\/strong><\/p> <p>[1] Abdollahinejad, Golnoosh, Saleh Baghersalimi, Denisa-Andreea Constantinescu, Sergey Shevchik, and David Atienza. &ldquo;<strong>AquaCast<\/strong>: Urban Water Dynamics Forecasting with Precipitation-Informed Multi-Input Transformer.&rdquo; <a href='https:\/\/arxiv.org\/abs\/2509.09458'><em>https:\/\/arxiv.org\/abs\/2509.09458<\/em><\/a><\/p> <p>[2] UrbanTwin project: <a href='https:\/\/urbantwin.ch\/'>https:\/\/urbantwin.ch\/<\/a>&nbsp;<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br> Contact email: <a href='mailto:golnoosh.abdollahinejad@epfl.ch; denisa.constantinescu@epfl.ch; david.atienza@epfl.ch?subject=When Sensors Go Silent: Masked Multivariate Transformers for Robust Urban Water Forecasting'>golnoosh.abdollahinejad@epfl.ch; denisa.constantinescu@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project708minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br>\";<\/script>\n<span id=project708><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br> <a href=#_ onclick=opendesc('project708',project708); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/urbantwin.ch target=_blank title='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/197.png width=70 alt='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor691><\/a><b><span style='font-size: 20px;'>Fast compilation for Coarse-Grained Reconfigurable Arrays (CGRAs) via Graph Minors<\/b><br><script>var project691=\"<p>This project offers students the opportunity to develop an efficient  mapping algorithm for deploying applications on Coarse-Grained  Reconfigurable Array (CGRA) accelerators [1]. Its primary goal is to  create a fast and scalable mapping methodology that achieves minimal  energy consumption and reduced runtime.<\/p>    <p>CGRA is a 2-dimensional mesh of Processing Elements (PEs) supporting  the execution of arithmetic operations (e.g., add, sub, mul), where each  PE comprises a local register file and an Arithmetic Logic Unit (ALU).  By being programmable at the granularity of operations, CGRAs are very  promising because they offer an in-between solution between the  efficiency of fixed-function hardware accelerators and the versatility  of bit-programmable FPGAs. However, they require substantial effort to  deploy applications, as operations must be mapped both in the space and  time dimensions.<\/p>    <p>Hence, most existing mapping methods tackle this challenge with a  limited scope, usually focusing on mapping a single loop onto the  hardware. While some approaches, such as the one proposed in [2], try to  extend support for control operations (e.g., if statements and\/or while  loops), they depend on integer linear programming models that do not  scale well, thus showing exponential growth in compilation time as the  application size increases.<\/p>    <p>In contrast, graph-based methods [3] have shown promise for  significantly faster mapping. This project aims to adopt this paradigm  for implementing a graph-minor-based mapping algorithm supporting to a  wide range of applications, including those with complex control  structures. The algorithm will automatically generate assembly code from  diverse workloads, which will be evaluated to assess its performance in  terms of latency and energy efficiency.<\/p>    <p><strong>Tasks description<\/strong><\/p>    <ol class='wp-block-list'><li>Understand the architecture of CGRAs, including their components, interconnects, and execution model.<\/li><li>Familiarize with the existing compilation flow developed in ESL-EPFL (Compigra).<\/li><li>Understand the graph minor problems in graph theory.<\/li><li>Implement an Operation Mapping Algorithm base on a graph minor approach.<\/li><li>Evaluate the application performance on CGRA with existing benchmarks [4].<\/li><\/ol>    <p><strong>Project objectives<\/strong><\/p>    <ul class='wp-block-list'><li>[Mandatory] Implement a dataflow graph mapping algorithm based on graph minors, targeting the OpenEdge CGRA platform [1].<\/li><li>[Optional] Extend the mapping capability to support control-flow graphs by incorporating graph minor detection prerequisites.<\/li><li>[Optional] Develop graph transformation techniques on subgraph manipulation when direct graph minor mapping fails.<\/li><li>Successful completion of the optional objectives may lead to a scientific publication based on the project&rsquo;s outcomes.<\/li><\/ul>    <p><strong>Required knowledge and skills<\/strong><\/p>    <ul class='wp-block-list'><li>Strong programming skills in C\/C++.<\/li><li>Strong foundation in mathematics, particularly graph theory.<\/li><li>Basic understanding of computer architecture.<\/li><li>Familiarity with LLVM and\/or MLIR frameworks is a plus.<\/li><li>Scientific curiosity.<\/li><\/ul>    <p><strong>References<\/strong><\/p>    <p>[1] &Aacute;lvarez, R. R., Denkinger, B., Sapriza, J., Calero, J. M.,  Ansaloni, G., &amp; Alonso, D. A. (2023, May). An open-hardware  coarse-grained reconfigurable array for edge computing. In <em>Proceedings of the 20th ACM International Conference on Computing Frontiers<\/em> (pp. 391-392).<\/p>    <p>[2] Wang, Y., Tirelli, C., Orlandic, L., Sapriza, J., &Aacute;lvarez, R. R.,  Ansaloni, G., &amp; Atienza, D. (2025, April). An mlir-based  compilation framework for cgra application deployment. In&nbsp;<em>International Symposium on Applied Reconfigurable Computing<\/em>&nbsp;(pp. 33-50). Cham: Springer Nature Switzerland.<\/p>    <p>[3] Zhou, G., Stojilovi\u0107, M. and Anderson, J. H., &ldquo;GRAMM: Fast CGRA  Application Mapping Based on A Heuristic for Finding Graph Minors,&rdquo; In <em>33rd International Conference on Field-Programmable Logic and Applications (FPL)<\/em>, Gothenburg, Sweden, 2023.<br \/><br \/>[4] M. Guthaus, J. Ringenberg, D. Ernst, T. Austin, T. Mudge, and R. Brown, &quot;Mibench: A free, commercially representative embedded benchmark suite,&quot; in Proceedings of the Fourth Annual IEEE International Workshop on Workload Characterization, 2001.<\/p>    <p><strong>Type of work<\/strong><\/p>    <ul class='wp-block-list'><li>80% <strong>SW design<\/strong>: development of the mapping algorithm that automate the generation of the assembly code.<\/li><li>20% <strong>benchmarking<\/strong> of the performance of generated assembly code.<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> Contact email: <a href='mailto:yuxuan.wang@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch?subject=Fast compilation for Coarse-Grained Reconfigurable Arrays (CGRAs) via Graph Minors'>yuxuan.wang@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project691minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project691><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project691',project691); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor655><\/a><b><span style='font-size: 20px;'>Microarchitectural explorations for biomedical applications<\/b><br><script>var project655=\"<p>Microcontrollers (MCUs) are used in a wide range of applications, from wearable devices and sensor monitoring, to robotics and automotive. In particular, the design of low-power microcontrollers for wearables in the biomedical domain has received a lot of attention in recent decades. Recent proposals such as BiomedBench [1] have created a collection of biomedical applications and kernels with the aim of informing the design of new processing architectures for wearable devices.<\/p> <p>In particular, the Embedded Systems Laboratory is developing <a href='https:\/\/x-heep.epfl.ch'>X-HEEP<\/a>, (eXtendable Heterogeneous Energy-Efficient Platform), which is an open-source, configurable, and extensible single-core RISC-V 32b MCU, sponsored by the <a href='https:\/\/ecocloud.epfl.ch'>EcoCloud<\/a> sustainable computing center of EPFL. X-HEEP is based on third-party open-source IPs and in-house IPs developed at ESL jointly with other EPFL laboratories. X-HEEP provides a framework to run applications compiled for RISC-V on a simulator (Verilator, Questasim, or VCS), on a Xilinx FPGA, and can be implemented in silicon as well. The first ASIC based on X-HEEP is called HEEPocrates.<\/p> <p><a href='https:\/\/biomedbench.epfl.ch'>BiomedBench <\/a>has recently been ported to X-HEEP. The open source nature of the platform, and the fact that it is being developed at ESL, creates an excellent chance to investigate which architectural features of low-power microcontrolers can increase the energy efficiency of wearables in the biomedical domain.<\/p> <p>In this project we want to explore if the use of an in-order superscalar core, in place of the in-order single-scalar core currently used in X-HEEP, the RISC-V OpenHW Group CV32E20 [2], can improve energy efficiency during the processing phase of the applications in BiomedBench. The working hypothesis will be that, whereas out-of-order execution introduces too much power overhead in comparison with the improvements obtained in execution time, in-order superscalar execution, in particular with the RISC-V ISA, produces larger time improvements than power increases. Therefore, in-order superscalar execution is suitable for reducing energy consumption in RISC-V microprocessors for the biomedical wearable domain. As a second step, we will evaluate if other architectural extensions, e.g., specifically for fixed-point arithmetic, can improve the energy efficiency of the microcontroller.<\/p> <p>During this project, the student will first develop a simulator for the RISC-V architecture that supports execution of the applications ported to X-HEEP (binary compatibility). Instead of developing a simulator from scratch, it will also be possible to use other existing open-source solutions, as long as they can be used for the second phase. The simulator will be used to generate dynamic execution traces from the applications in BiomedBench.<\/p> <p>In a second phase, the student will modify the simulator to evaluate the impact on performance of in-order superscalar execution. The simulator will be easily modifiable, so that it will be possible to test which combinations of additional functional units will produce the best performance impact with the minimal cost in HW.<\/p> <p>In a third phase, the student will evaluate, with help from the people participating in the X- HEEP project, the impact on power of the best candidate architectures found based on performance improvement. In this way, at the end of the project it will be possible to determine which architectural optimizations produce the best improvements in terms of total energy consumption, based on the maximum performance improvement with the minimum additional power.<\/p> <p>The previous explorations do not require modifying the traces obtained from the execution in the simulator. Therefore, the compiled binary used for X-HEEP will be valid during these phases. However, other explorations, such as the introduction of specific instructions for fixed-point execution, may need modifying the execution trace. This will be done either using the original assembly code produced by the compiler, or dynamically modifying the trace during simulation.<\/p> <p>The expected outcomes of this project are:<\/p> <ul id='l1'> <li> <p>Development of a lightweight RISC-V simulator that can execute the binary (compiled) applications of BiomedBench for X-HEEP. No interrupts will be included in the simulations. The simulator will produce as output a complete memory dump and a summary of processor cycles required for execution of the benchmark.<\/p> <ul id='l2'> <li data-list-text='\u25cb'> <p>The correctness of the simulation will be guaranteed at all times comparing the output of the application with the expected outputs from BiomedBench.<\/p> <\/li> <\/ul> <\/li> <li> <p>Generation of dynamic execution traces from the BiomedBench applications using the simulator. The student will be allowed to propose a different mechanism to obtain the traces, as long as it allows them to conduct the explorations in the following phases.<\/p> <\/li> <li> <p>Modification of the simulator to account for superscalar execution of the traces. At this point, no extra functional units will be introduced; the simulation will only account for data dependencies between the instructions, the nature of the instructions (e.g., whether the first instruction is a branch or not), and the availability of resource classes. For example, at this stage: loads can proceed in parallel with additions; one addition and one multiplication can be executed simultaneously; two additions\/subtractions cannot be executed in parallel.<\/p> <p style='padding-left: 5pt;text-indent: 0pt;text-align: left'>Optional\/additional outcomes:<\/p> <\/li> <li> <p>Exploration of the performance benefits of introducing different types of arithmetic operators, e.g., dual adders.<\/p> <\/li> <li> <p>Exploration of the overhead in area and power of the proposed modifications of the control stage and additional functional units.<\/p> <\/li> <li> <p>Exploration of other optimizations specific to the biomedical domain (e.g., for fixed- point arithmetic), as driven by the BiomedBench applications.<\/p> <p>Throughout the project, the student will learn:<\/p> <\/li> <li> <p>Basic processor architecture concepts and the RISC-V ISA.<\/p> <\/li> <li> <p>The main features of applications in the biomedical wearable domain.<\/p> <\/li> <li> <p>Advanced processor architecture concepts such as superscalar, in-order and out-of- order execution.<\/p> <\/li> <li> <p>How to work with git repositories in a team of contributors to the same project.<\/p> <\/li> <\/ul> <p>The project will be carried out at the ESL at EPFL, one of the world\u2019s top-class universities, including EcoCloud\u2019s technical support. ESL is an active group <span style='color: #000'>(24 Ph.D. students among 45 members) <\/span>involved in many research lines. The student will be under the supervision of Prof. David Atienza (ESL) and Dr. Miguel Pe\u00f3n-Quir\u00f3s (EcoCloud), with technical support from Stefano Albini (ESL).<\/p> <h4>Project objectives:<\/h4> <ol id='l3'> <li data-list-text='1.'> <p>Understanding the RISC-V architecture and development of a lightweight simulator that can execute the applications of BiomedBench compiled for X-HEEP and produce execution statistics.<\/p> <\/li> <li data-list-text='2.'> <p>Modification of the simulator to evaluate in-order superscalar (parallel) execution of the applications in BiomedBench, using the original or additional numbers of functional units. The output of this evaluation will sustain or refute our initial hypothesis.<\/p> <\/li> <li data-list-text='3.'> <p>(Optional) Evaluation of the overheads in terms of area\/power of the proposed modifications.<\/p> <\/li> <li>(Optional) Proposal of additional architectural improvements specific for the domain of biomedical wearables.<\/li><p><\/p> <\/ul><h4>Required knowledge and skills:<\/h4> <ul id='l4'> <li> <p>C++ and Python. General Linux use and scripting.<\/p> <\/li> <li> <p>Good background in computer architecture and algorithms.<\/p> <\/li> <li> <p>Some familiarity with any assembly language (RISC-V ISA will be used throught the project).<\/p> <\/li> <li> <p>Good analytical skills.<\/p> <\/li> <li> <p>Teamwork and git.<\/p><\/li><\/ul> <h4>Appreciated skills:<\/h4> <ul><li> <p>Scientific curiosity.<\/p> <\/li> <li> <p>Good communication skills.<\/p> <\/li> <li> <p>Advanced English (interaction during the project will be in English).<\/p> <\/li> <\/ul>  <\/ol> <h4>Type of work: <span class='p'>40% theory analysis, 60% design and simulation.<\/span><\/h4> <p class='s7' style='padding-top: 14pt;padding-left: 5pt;text-indent: 0pt'>[1]. Dimitrios Samakovlis et al. \u201cBiomedBench: A benchmark suite of TinyML biomedical applications for low- power wearables.\u201d IEEE Design &amp; Test, 2024. <a href='https:\/\/infoscience.epfl.ch\/handle\/20.500.14299\/208450.7'>https:\/\/infoscience.epfl.ch\/handle\/20.500.14299\/208450.7<\/a><\/p> <p class='s8'>[2]. Pasquale Davide Schiavone et al. \u201cSlow and steady wins the race? A comparison of ultra-low-power RISC-V cores for Internet-of-Things applications\u201d. In: Int. Symp. on Power and Timing Modeling, Optimization and Simulation (PATMOS). <a href='https:\/\/ieeexplore.ieee.org\/document\/8106976'>IEEE. 2017, pp. 1\u20138.<\/a><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Miguel Pe\u00f3n-Quir\u00f3s, EcoCloud; Prof. David Atienza, ESL; Stefano Albini, ESL <br> Contact email: <a href='mailto:miguel.peon@epfl.ch; david.atienza@epfl.ch; stefano.albini@epfl.ch?subject=Microarchitectural explorations for biomedical applications'>miguel.peon@epfl.ch; david.atienza@epfl.ch; stefano.albini@epfl.ch<\/a><br>\";<\/script>\n<script>var project655minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Miguel Pe\u00f3n-Quir\u00f3s, EcoCloud; Prof. David Atienza, ESL; Stefano Albini, ESL <br>\";<\/script>\n<span id=project655><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Miguel Pe\u00f3n-Quir\u00f3s, EcoCloud; Prof. David Atienza, ESL; Stefano Albini, ESL <br> <a href=#_ onclick=opendesc('project655',project655); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/wp-content\/uploads\/2024\/11\/Microarchitectural-explorations-for-biomedical-applications.pdf title='document link'><img width=30 border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/doclink.gif hspace=2 alt='document link'><\/a><\/td><td><a href='https:\/\/biomedbench.epfl.ch\/' title='weblink'><img hspace=2 width=30 border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/weblink.gif alt='web link'><\/a><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor637><\/a><b><span style='font-size: 20px;'>Enabling Local Tightly, Global Loosely coupled Programmable Accelerators  Heterogeneous Systems by integrating X-HEEP into ESP SoCs<\/b><br><script>var project637=\"<p><i>Artificial Intelligence (AI) <\/i>has been one of the most dominant factors driving technology innovation over the last decade and has been exploited in a huge variety of fields ranging from image recognition and natural language processing to autonomous driving and modeling of complex physical systems. For low-power applications, software-programmable microcontrollers (MCUs) are preferred thanks to their versatility and short time-to-market. However, MCUs are often extended with domain-specific accelerators to meet timing and energy requirements. Examples of applications that benefit from on-device AI include small cameras that recognize faces, microphones that filter background noise and recognize voice commands, and wearable Galvanic Skin Response sensors that detect emotions, etc. Most of such applications are implemented using neural networks, deployed partially or completely at the edge, on tiny and ultra-low-power SoCs. Many domain-specific accelerators have been proposed to speed up the computations of such kernels, such as feature-extraction engines, matrix-multiplication, or convolutional accelerators.&nbsp; SoCs usually employ one or more accelerators, giving birth to heterogeneous systems.&nbsp;<\/p> <p>Such accelerators are called \u201ctightly coupled\u201d when they are integrated close to the CPU so that the synchronization and data exchange between them are shallow and require low latency. In such systems, the accelerators and CPUs are usually connected on the same bus.<\/p> <p>Otherwise, they are called \u201cloosely coupled.\u201d In these systems, the accelerators and CPUs are not connected in the same bus but rather in a network-on-chip (NoC). This allows for more scalable SoCs but at the price of slower and more complex communications.<\/p> <p>&nbsp;<\/p> <p>X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform) is an open-source, configurable, and extensible single-core RISC-V 32-bit MCU developed at the Embedded Systems Laboratory (ESL), sponsored by the EcoCloud Sustainable Computing Center of Swiss Federal Institute of Technology Lausanne (EPFL).<\/p> <p>It has been designed based on existing open-source IPs from the PULP, OpenHW Group, and OpenTitan projects. Its main advantage is that it eases the integration of tightly coupled accelerators throughout the so-called eXtension-Accelerator interface (XAIF) into SoCs. So far, it has been implemented with accelerators such as CGRAs, Near-Memory Computing IPs, Systolic Arrays, POSIT datapaths, and GPUs.<\/p> <p>ESP (Embedded Scalable Platform) is an open-source platform for heterogeneous SoC design that provides a flexible tile-based architecture built on a multi-plane NoC. It was developed at Columbia University by the System-Level Design (SLD) group and is also compatible with RISC-V IPs. ESP also provides a flow to develop and integrate accelerators described in High-Level Synthesis (SystemC, C++) or RTL (Verilog\/SystemVerilog\/VHDL) loosely coupled via the NoC with a specified tile-based interface.&nbsp;<\/p> <p>In this project, we propose to build a Local-Tightly, Global-Loosely coupled Accelerator Heterogeneous System by integrating X-HEEP into an ESP SoC.&nbsp;<\/p> <p>We want to use ESP to build the main tile-based SoC and X-HEEP to build the internals of individual accelerator tiles.&nbsp;<\/p> <p>This highly programmable and heterogeneous SoC will have several levels of accelerator integration. In fact, inside a tile, the accelerator will be seen as tightly coupled by the local CPU within the tile (X-HEEP), while it will be seen as loosely coupled by the CPU(s) integrated into another tile.<\/p> <p>The final system can leverage programmable tiles based on flexible CPU+Accelerator architectures, where each tile is specialized for a given function. For example, an ESP system could be composed of one tile specialized in running Linux with a RISC-V CPU; one tile could integrate X-HEEP with a GPU for high-parallel functions; one tile could contain X-HEEP with a CGRA for highly spatial reconfigurable functions; and one final tile with X-HEEP with near-memory computing IPs to efficient local processing and last-level cache functions. Plus, extra tiles for memories and I\/O.<\/p> <p><b>Project objectives:<\/b><\/p> <p>Project objectives:<\/p> <ol> <li>Understanding the ESP SoC, with a particular focus on the tile interface and NoC protocol. A provided example should be run and emulated on the FPGA or simulated with Questasim.<\/li> <li>Understanding the X-HEEP SoC, with a particular focus on the XAIF interface and bus protocol. A provided example should be run and emulated on the FPGA or simulated with Questasim.<\/li> <li>Build an interface and its relative bridge to connect X-HEEP and ESP and integrate X-HEEP into ESP in a minimalistic configuration to verify and test an application on Questasim or FPGA. To improve compatibility test coverage, one tile must be X-HEEP, while the other must be a tile coming from the ESP IPs. Such a test should include a C program that enables the bidirectional communication of the two tiles.<\/li> <li>Build an ESP SoC based on an ESP\u2019s example capable of running Linux with the CVA6, including an X-HEEP tile enhanced with a tightly coupled accelerator. The final system should leverage X-HEEP to accelerate one AI function.<\/li> <\/ol> <p>Throughout the project, the student will learn:<\/p> <ul> <li>how to design, use, and leverage complex SoCs such as ESP and X-HEEP.<\/li> <li>how to integrate and test different systems, simulating and emulating them on Questasim and\/or FPGA.<\/li> <li>how to design interfaces and specifications.<\/li> <li>how to deploy an application on a complex heterogeneous RISC-V SoC microcontroller.<\/li> <li>how to work with version control (Git) and third-party, open-source repositories.<\/li> <li>How to work in a collaborative team of people from different universities, all contributing to the same and different projects.<\/li> <\/ul> <p>The project will be carried out at the ESL at EPFL, one of the world\u2019s top-class universities. ESL is an active group (24 Ph.D. students among 45 members) involved in many research aspects, therefore providing a stimulating research environment.&nbsp;<\/p> <p>The student will be supervised by Prof. David Atienza, Dr. Davide Schiavone, Prof. Daniele Jahier Pagliari, and Prof. Alessio Burrello from the Polytechnic of Turin.<\/p> <p>In addition, given an ongoing collaboration with SDL Group at Columbia<\/p> <p>University that has developed ESP, Prof. Luca Carloni from Columbia University might also become a co-supervisor.<\/p> <p><b>Required knowledge and skills:<\/b><\/p> <ul> <li>RTL design and FPGA implementation in SystemVerilog<\/li> <li>Good understanding of memory architectures and microcontrollers<\/li> <li>Good analytical skills<\/li> <li>Good background in computer architecture<\/li> <li>Teamwork and git<\/li> <\/ul> <p><b>Appreciated skills:<\/b><\/p> <ul> <li>Scientific curiosity<\/li> <li>Good communication skills<\/li> <li>Advanced English <\/li> <\/ul> <p><b>Type of work:<\/b> 40% theory analysis, 60% SW\/HW co-design and simulation<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Prof. Daniele Jahier Pagliari, Prof. Alessio Burrello, Prof. David Atienza<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch;daniele.jahier@polito.it;alessio.burrello@polito.it;david.atienza@epfl.ch?subject=Enabling Local Tightly, Global Loosely coupled Programmable Accelerators  Heterogeneous Systems by integrating X-HEEP into ESP SoCs'>davide.schiavone@epfl.ch;daniele.jahier@polito.it;alessio.burrello@polito.it;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project637minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Prof. Daniele Jahier Pagliari, Prof. Alessio Burrello, Prof. David Atienza<br>\";<\/script>\n<span id=project637><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Prof. Daniele Jahier Pagliari, Prof. Alessio Burrello, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project637',project637); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor565><\/a><b><span style='font-size: 20px;'>Porting and Optimization of biomedical benchmark applications to the X-HEEP Open-Source RISC-V-based microcontroller platform<\/b><br><script>var project565=\"<p>Wearable devices promise to improve  preventive medicine through continuous health monitoring of chronic  diseases. The design of low-power wearables for the biomedical domain  has received much attention in recent decades, as technological  advancements in chip manufacturing have allowed real-time monitoring of  patients within the &micro;W range. To ensure continued progression in this  domain, a co-design view that optimizes both hardware and software  simultaneously, and standardized tools are necessary.<\/p>    <p>In response to the aforementioned needs, the Embedded Systems  Laboratory(ESL) has developed BiomedBench. BiomedBench is a new  benchmark suite composed of state-of-the-art (SoA) biomedical  applications for real-time monitoring of patients using wearable  devices. Each application presents different requirements during the  typical signal acquisition and processing phases, including varying  computational workloads and relations between active and idle times.  Therefore, BiomedBench provides hardware developers with a tool to  assess the efficiency of their ultra-low power (ULP) platform designs  under varying requirements. Moreover, the open-sourcing nature of  BiomedBench will serve as a baseline for future application developers  aspiring to develop and deploy their biomedical applications in ULP  devices.<\/p>    <p>Typically, biomedical applications for patient monitoring tasks  include the modules depicted in Figure 1 below. Typically, the  processing step consists of signal preprocessing, feature extraction,  and inference based on these features. However, applications can exhibit  a wide range of workloads and computational requirements. For example,  feature extraction can be implemented explicitly (that is, manually  engineered features) or implicitly (e.g., convolutional neural network  (CNN)). Similarly, the inference step can use a lightweight machine  learning method, such as a random forest or a complex deep neural  network (DNN).<\/p>    <img width=500 src='https:\/\/ecocloud.ch\/extranet\/?p=200&amp;preview=true' alt='' \/>    <p><img src='https:\/\/ecocloud.ch\/extranet\/wp-content\/uploads\/2023\/12\/figure1.png' alt='' \/><br \/><strong>Figure 1: Typical computational pipeline of biomedical applications<\/strong><\/p>    <p>From an implementation point of view, the system undergoes an  always-on acquisition phase and an intermittent processing phase, as  presented in Figure 2. A complete processing period consists of an idle  period, during which the processing unit is in low-power mode, and a  computation upon acquisition of the full input signal. The duration of  the idle period varies significantly between applications and can  dominate the system&rsquo;s energy consumption.<\/p>    <p><img src='https:\/\/ecocloud.ch\/extranet\/wp-content\/uploads\/2023\/12\/figure2.png' alt='' width='496' height='137' \/><br \/><strong>Figure 2: System operating modes during signal acquisition<\/strong><\/p>    <p>BiomedBench includes eight applications representative of the  biomedical domain that offer a variety of workloads and profiles for the  processing, idle, and acquisition phases. All applications are coded in  C or C++. Four applications are implemented in fixed-point arithmetic,  targeting low-end MCUs. The rest are implemented in 32-bit floating  point arithmetic. Four applications also include a multi-core  implementation that enables significant acceleration in the presence of  multiple cores.<\/p>    <h2 class='wp-block-heading'>Considered ULP hardware platform &ndash; X-HEEP \/ HEEPocrates<\/h2>    <p>On the hardware side, ESL has devoted a lot of research efforts to  developing a new open-source hardware platform, called X-HEEP  (eXtendable Heterogeneous Energy-Efficient Platform), to support the  monitoring of participants in clinical studies with low energy  footprint. X-HEEP is an open-source, configurable, and extensible  single-core RISC-V 32b MCU, sponsored by the EcoCloud Sustainable  Computing Center of EPFL. It is based on many third-party open-source  IPs and in-house IPs developed at the Embedded Systems Laboratory (ESL)  jointly with other EPFL laboratories. X-HEEP provides a framework to run  applications compiled for RISC-V on<\/p>    <p>a simulator (Verilator, Questasim, or VCS), on a Xilinx FPGA, and can be implemented in silicon as well.<\/p>    <p>In 2023, ESL fabricated HEEPocrates, the first ASIC implementation  (in TSMC 65nm) deploying X-HEEP configured with the cv32e2 core and with  256kB of memory. HEEPocrates instantiates X-HEEP as the main  microcontroller driving a CGRA, an In-Memory Computing macro.  HEEPocrates belongs to the category of ULP platforms featuring a 6mm2  X-HEEP chip, a maximum frequency of 470 MHz consuming up to 48mW. Hence,  HEEPocrates is a suitable platform to deploy the BiomedBench  applications on and conduct a performance and energy analysis.<\/p>    <h2 class='wp-block-heading'>Thesis summary<\/h2>    <p>The goal of this thesis is to <strong>utilize BiomedBench to evaluate the X-HEEP and HEEPocrates platforms<\/strong>. To achieve this, the student will have to:<\/p>    <ol><li><strong>Learn to deploy the BiomedBench applications on X-HEEP \/ HEEPocrates <\/strong>by efficiently utilizing the capabilities of the platform for the sleep and acquisition phases<\/li><li><strong>Perform timing and energy measurements <\/strong>of each application running on X-HEEP \/HEEPocrates, analyze and compare with SoA results.<\/li><li><strong>Deploy the BiomedBench applications on X-HEEP FPGA <\/strong>changing  the memory size and CPU to find the optimal X-HEEP configuration for  BiomedBench, including data transfers from the FLASH to the on-chip SRAM  when data overfit the internal capacity.<\/li><li>(OPTIONAL) <strong>Apply algorithmic and software optimizations <\/strong>on  each application to speed up computations or\/and reduce memory  footprint in HEEPocrates without degradation of the final  application-level accuracy result.<\/li><\/ol>    <h2 class='wp-block-heading'>Thesis outcome<\/h2>    <p>The outcome of the M.Sc. thesis will be published open-source in the BiomedBench and X-HEEP (<a href='https:\/\/github.com\/esl-epfl\/x-heep'>link<\/a>) repositories: The expected outcomes of this thesis are:<\/p>    <ul><li>[Deployment of complete applications to X-HEEP and\/or HEEPocrates]  Development of software to run all the applications on X-HEEP \/  HEEPocrates. The basic C\/C++ implementation of each application is  given, but the porting to the platforms requires some extra code and  smart deployment decisions to respect the memory constraints of the  platform. Complementary to the processing part of each application, the  acquisition and sleep mode should be programmed efficiently.<\/li><li>[Performance and energy results] Measuring performance and energy  consumption running each complete application on X-HEEP \/ HEEPocrate and  comparing with other state-of-the-art platforms.<\/li><li>[Finding the  optimal] Find the optimal configuration of X-HEEP FPGA implementation  for the benchmark by varying the CPU and memory capacity.<\/li><li>[Application optimization] Optimize the C\/C++ implementation and\/or  the algorithms involved in each application. The target of the  optimizations is to improve the energy efficiency of the complete  application which can be achieved through decreasing the execution time  or the memory footprint, provided there is no accuracy degradation in  the final result.<\/li><\/ul>    <h2 class='wp-block-heading'>Learning outcome<\/h2>    <p>Throughout the thesis, the student will learn:<\/p>    <ul><li>How different real-time patient monitoring applications are structured and what is the state-of-the-art in the domain<\/li><li>How to deploy such applications in resource-constrained devices<\/li><\/ul><ul><li>How to orchestrate the processing, acquisition, and sleeping phases of such applications<\/li><\/ul><ul><li>How to identify application bottlenecks and apply algorithmic or software optimizations<\/li><\/ul><ul><li>How  to study the critical architectural features of a platform, such as  X-HEEP, to achieve maximum energy efficiency for each application  deployment<\/li><\/ul><ul><li>How to use Git to manage projects with multiple developers<\/li><li>How to collaborate with the team and to analyze and present the results<\/li><\/ul>    <p>The thesis will be carried out at the ESL at EPFL, one of the world&rsquo;s  top-class universities. ESL is an active group (24 Ph.D. students among  45 members) involved in many research aspects. The student will be  under the supervision of Prof. David Atienza, Dr. Davide Schiavone, and  two Ph.D. students (Dimitrios Samakovlis and Stefano Albini).<\/p>    <h2 class='wp-block-heading'>Required knowledge and skills:<\/h2>    <ul><li>Low-level software design (C and\/or C++ is going to be used throughout the thesis)<\/li><li>Good understanding of memory architectures and microcontrollers<\/li><li>Good analytical skills<\/li><li>Makefiles for complex project structures<\/li><li>Teamwork and git<\/li><li>Good background in algorithms and common ML models (required for the  optimization part at the end of the project, which is optional)<\/li><\/ul>    <h2 class='wp-block-heading'>Appreciated skills:<\/h2>    <ul><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English<\/li><li>Assembly knowledge (useful for in-depth performance analysis)<\/li><\/ul>    <p><strong>Type of work: <\/strong>10% theory analysis, 90% coding and experimenting<\/p>                                                                                              <h4 class='section-title'><br \/><\/h4><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Mr. Dimitrios Samakovlis, Mr. Stefano Albini, Prof. David Atienza<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch;dimitrios.samakovlis@epfl.ch;stefano.albini@epfl.ch; david.atienza@epfl.ch;?subject=Porting and Optimization of biomedical benchmark applications to the X-HEEP Open-Source RISC-V-based microcontroller platform'>davide.schiavone@epfl.ch;dimitrios.samakovlis@epfl.ch;stefano.albini@epfl.ch; david.atienza@epfl.ch;<\/a><br>\";<\/script>\n<script>var project565minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Mr. Dimitrios Samakovlis, Mr. Stefano Albini, Prof. David Atienza<br>\";<\/script>\n<span id=project565><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Mr. Dimitrios Samakovlis, Mr. Stefano Albini, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project565',project565); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=https:\/\/ecocloud.ch\/extranet\/wp-content\/uploads\/2023\/12\/BiomedBench_X-HEEP-Master-Thesis-Proposal-Spring-2024.pdf title='document link'><img width=30 border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/doclink.gif hspace=2 alt='document link'><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor557><\/a><b><span style='font-size: 20px;'>Implementation of an Accelerator based on Near-Memory Computing IPs for a RISC-V-based Microcontroller<\/b><br><script>var project557=\"<p>Artificial Intelligence (AI) has been one of the most dominant factors driving technology innovation over the last decade and has been exploited in a huge variety of fields ranging from image recognition and natural language processing to autonomous driving and modeling of complex physical systems. As integrated semiconductor devices become smaller and faster, systems-on-chip (SoCs) become more and more complex, enabling pocket-size, wearable, battery-powered systems to efficiently support the computationally expensive algorithms at the core of complex AI models while running on a limited power budget. Due to the large number of parameters, software-programmable SoCs are preferred thanks to their versatility and short time-to-market. Small cameras that recognize faces, microphones that filter background noise or recognize voice commands, wearable devices that detect epilepsy attacks, and implantable devices that constantly monitor tens of body parameters and release drugs accordingly to prevent organ failures are just a few examples of what smart edge devices represent in the AI revolution we are experiencing.<\/p> <p>One of the main bottlenecks for performance and energy efficiency of next-generation SoCs reside in the limited memory bandwidth inherent in the traditional Von Neuman architecture. One idea to overcome this limitation is to bring computation within the memory subsystem, to better exploit the available memory bandwidth and leverage data reuse more efficiently. Such computational paradigm is referred to as Processing-in-Memory (PiM) or Compute-in-Memory (CiM). Further benefits can be achieved by leveraging the Single Instruction, Multiple Data (SIMD) approach, where the same operation (i.e., an instruction) operates on a multitude of data (e.g., a vector or matrix), therefore significantly reducing the number of instructions loaded from memory and contributing to reducing the system\u2019s energy consumption.<\/p> <p>The <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded Systems Laboratory (ESL)<\/a> at the Swiss Federal Institute of Technology Lausanne (EPFL) has developed two SRAM-based low-power architectures (known as Caesar and Carus) that normally behave as traditional memories, but also offer scalar and vector computing capabilities (i.e., arithmetic and logic operations such as addition, and, or, xor, multiplication, multiply-add, etc.) between two or more memory words. Because the data is processed within the memory layout itself, these smart near-memory IPs eliminate the need for moving operands through the system bus and into the local memory elements of processing elements that are physically far from the memory (e.g., inside the system CPU).&nbsp;<\/p> <p>In addition to memory architectures, ESL has also developed X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform). It is an open-source, configurable, and extensible single-core RISC-V 32-bit Microcontroller Unit (MCU), sponsored by the EcoCloud Sustainable Computing center of EPFL. It is based on many third-party open-source IPs as well as in-house IPs developed at the ESL jointly with other EPFL laboratories. X-HEEP provides a framework to configure and extend the MCU and experiment with it as an RTL simulation model (Verilator, Questasim, or VCS), a hardware prototype on a Xilinx FPGA, and even tape it out as a standalone ASIC circuit. The framework also provides the RISC-V software toolchain and the SDK that are necessary to deploy applications on the MCU.<\/p> <p>Recently, the X-HEEP system has been extended to integrate Caesar- and Carus-based memories besides traditional SRAM banks. When running in computing mode, they can be programmed or controlled using dedicated software routines that implement application-specific computing kernels (e.g., matrix multiplication). Otherwise, they operate as traditional memories. By definition, these near-memory computing units exclusively process data that is directly mapped inside their private memory space (i.e., the memory banks instantiated within the IP itself.<\/p> <p>From a low-level point of view,&nbsp; this approach reduces data movement and memory bandwidth, thus increasing the system\u2019s energy efficiency. However,&nbsp; from an application point of view, it limits the size of the data that can be processed by the in-memory computing kernel (as it must fit inside a single memory bank) and does not allow for multi-memory bank parallelism opportunities.<\/p> <p>Many edge AI applications rely on fixed-point operations, replacing the more expensive floating-point operations used when deploying the same machine learning models on more powerful hardware. As of today, Carus and Caesar support integer datapath on generic 32, 16, and 8-bit instructions. None of the operations is specifically designed to deal with the fixed-point data format (as additions or multiplications followed by rounding and shifting instructions). Therefore, fixed-width operation must be emulated in software in the current implementation.<\/p> <p>This thesis aims to extend the Instruction Set Architecture (ISA) of the Carus near-memory computing IP with fixed-point instructions to increase performance and energy efficiency.<\/p> <p>Throughout the project, the student will learn:<\/p> <ol class='wp-block-list'> <li>How the Carus NMC IP works and how to offload computationally expensive tasks to it within the X-HEEP framework.<\/li> <li>How to extend the Carus NMC IP decoder and execution pipeline to support fixed-point instructions as additions, subtractions, and multiplications with rounding and shift in 32, 16, and 8-bit modes.<\/li> <li>Verify the functionality of the new instructions with randomized inputs.<\/li> <li>Verify that the introduced modifications do not alter the timing characteristics of the system, and iterate on the architecture in case they do (for example with techniques such as multicycle logic paths).<\/li> <li><em>[Optional]<\/em> Update a few existing applications to use the new fixed-point instructions instructions and test them on the system deployed on an FPGA.<\/li> <li>How to work with version control (Git) and third-party, open-source repositories.<\/li> <li>How to work in a team of people all contributing to the same project.<\/li> <\/ol> <p>The project will be carried out at the ESL at EPFL, one of the world\u2019s top-class universities. ESL is an active group (24 PhD students among 45 members) involved in many research aspects, therefore providing a stimulating research environment. The student will be under the supervision of Prof. David Atienza, Dr. Davide Schiavone, and Dr. Michele Caon.<\/p> <p><strong>Project objectives:<\/strong><\/p> <p>Project objectives:<\/p> <ol class='wp-block-list'> <li>Design a new set of fixed-point instructions as addition, subtraction, multiplication, and multiply-add supporting 32-, 16-, and 8-bit data elements by extending the Carus NMC IP decoder and execution pipeline.<\/li> <li>Verify that such instructions work correctly with randomized tests by extending the Carus testbench.<\/li> <li>Verify that the timing characteristics (e.g., the maximum operating frequency) of Carus ASIC implementation do not get worse and that the area increases negligibly by checking its existing physical implementation flow. In case it does, iterate the hardware.<\/li> <li><em>[Optional]<\/em> Update existing applications to leverage the new instructions and run the application on the system\u2019s hardware model deployed on an FPGA.<\/li> <\/ol> <p><strong>Required knowledge and skills:<\/strong><\/p> <ul class='wp-block-list'> <li>RTL design and FPGA implementation in SystemVerilog<\/li> <li>Good understanding of memory architectures and microcontrollers<\/li> <li>Good analytical skills<\/li> <li>Good background in computer architecture<\/li> <li>Teamwork and Git<\/li> <\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul class='wp-block-list'> <li>Scientific curiosity<\/li> <li>Good communication skills<\/li> <li>Advanced English&nbsp;<\/li> <\/ul> <p><strong>Type of work:<\/strong> 40% theory analysis, 60% SW\/HW co-design and simulation<\/p><p>Artificial Intelligence (AI) has been one of the most dominant factors driving technology innovation over the last decade and has been exploited in a huge variety of fields ranging from image recognition and natural language processing to autonomous driving and modeling of complex physical systems. As integrated semiconductor devices become smaller and faster, systems-on-chip (SoCs) become more and more complex, enabling pocket-size, wearable, battery-powered systems to efficiently support the computationally expensive algorithms at the core of complex AI models while running on a limited power budget. Due to the large number of parameters, software-programmable SoCs are preferred thanks to their versatility and short time-to-market. Small cameras that recognize faces, microphones that filter background noise or recognize voice commands, wearable devices that detect epilepsy attacks, and implantable devices that constantly monitor tens of body parameters and release drugs accordingly to prevent organ failures are just a few examples of what smart edge devices represent in the AI revolution we are experiencing.<\/p> <p>One of the main bottlenecks for performance and energy efficiency of next-generation SoCs reside in the limited memory bandwidth inherent in the traditional Von Neuman architecture. One idea to overcome this limitation is to bring computation within the memory subsystem, to better exploit the available memory bandwidth and leverage data reuse more efficiently. Such computational paradigm is referred to as Processing-in-Memory (PiM) or Compute-in-Memory (CiM). Further benefits can be achieved by leveraging the Single Instruction, Multiple Data (SIMD) approach, where the same operation (i.e., an instruction) operates on a multitude of data (e.g., a vector or matrix), therefore significantly reducing the number of instructions loaded from memory and contributing to reducing the system\u2019s energy consumption.<\/p> <p>The <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded Systems Laboratory (ESL)<\/a> at the Swiss Federal Institute of Technology Lausanne (EPFL) has developed two SRAM-based low-power architectures (known as Caesar and Carus) that normally behave as traditional memories, but also offer scalar and vector computing capabilities (i.e., arithmetic and logic operations such as addition, and, or, xor, multiplication, multiply-add, etc.) between two or more memory words. Because the data is processed within the memory layout itself, these smart near-memory IPs eliminate the need for moving operands through the system bus and into the local memory elements of processing elements that are physically far from the memory (e.g., inside the system CPU).&nbsp;<\/p> <p>In addition to memory architectures, ESL has also developed X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform). It is an open-source, configurable, and extensible single-core RISC-V 32-bit Microcontroller Unit (MCU), sponsored by the EcoCloud Sustainable Computing center of EPFL. It is based on many third-party open-source IPs as well as in-house IPs developed at the ESL jointly with other EPFL laboratories. X-HEEP provides a framework to configure and extend the MCU and experiment with it as an RTL simulation model (Verilator, Questasim, or VCS), a hardware prototype on a Xilinx FPGA, and even tape it out as a standalone ASIC circuit. The framework also provides the RISC-V software toolchain and the SDK that are necessary to deploy applications on the MCU.<\/p> <p>Recently, the X-HEEP system has been extended to integrate Caesar- and Carus-based memories besides traditional SRAM banks. When running in computing mode, they can be programmed or controlled using dedicated software routines that implement application-specific computing kernels (e.g., matrix multiplication). Otherwise, they operate as traditional memories. By definition, these near-memory computing units exclusively process data that is directly mapped inside their private memory space (i.e., the memory banks instantiated within the IP itself.<\/p> <p>From a low-level point of view,&nbsp; this approach reduces data movement and memory bandwidth, thus increasing the system\u2019s energy efficiency. However,&nbsp; from an application point of view, it limits the size of the data that can be processed by the in-memory computing kernel (as it must fit inside a single memory bank) and does not allow for multi-memory bank parallelism opportunities.<\/p> <p>Many edge AI applications rely on fixed-point operations, replacing the more expensive floating-point operations used when deploying the same machine learning models on more powerful hardware. As of today, Carus and Caesar support integer datapath on generic 32, 16, and 8-bit instructions. None of the operations is specifically designed to deal with the fixed-point data format (as additions or multiplications followed by rounding and shifting instructions). Therefore, fixed-width operation must be emulated in software in the current implementation.<\/p> <p>This thesis aims to extend the Instruction Set Architecture (ISA) of the Carus near-memory computing IP with fixed-point instructions to increase performance and energy efficiency.<\/p> <p>Throughout the project, the student will learn:<\/p> <ol class='wp-block-list'> <li>How the Carus NMC IP works and how to offload computationally expensive tasks to it within the X-HEEP framework.<\/li> <li>How to extend the Carus NMC IP decoder and execution pipeline to support fixed-point instructions as additions, subtractions, and multiplications with rounding and shift in 32, 16, and 8-bit modes.<\/li> <li>Verify the functionality of the new instructions with randomized inputs.<\/li> <li>Verify that the introduced modifications do not alter the timing characteristics of the system, and iterate on the architecture in case they do (for example with techniques such as multicycle logic paths).<\/li> <li><em>[Optional]<\/em> Update a few existing applications to use the new fixed-point instructions instructions and test them on the system deployed on an FPGA.<\/li> <li>How to work with version control (Git) and third-party, open-source repositories.<\/li> <li>How to work in a team of people all contributing to the same project.<\/li> <\/ol> <p>The project will be carried out at the ESL at EPFL, one of the world\u2019s top-class universities. ESL is an active group (24 PhD students among 45 members) involved in many research aspects, therefore providing a stimulating research environment. The student will be under the supervision of Prof. David Atienza, Dr. Davide Schiavone, and Dr. Michele Caon.<\/p> <p><strong>Project objectives:<\/strong><\/p> <p>Project objectives:<\/p> <ol class='wp-block-list'> <li>Design a new set of fixed-point instructions as addition, subtraction, multiplication, and multiply-add supporting 32-, 16-, and 8-bit data elements by extending the Carus NMC IP decoder and execution pipeline.<\/li> <li>Verify that such instructions work correctly with randomized tests by extending the Carus testbench.<\/li> <li>Verify that the timing characteristics (e.g., the maximum operating frequency) of Carus ASIC implementation do not get worse and that the area increases negligibly by checking its existing physical implementation flow. In case it does, iterate the hardware.<\/li> <li><em>[Optional]<\/em> Update existing applications to leverage the new instructions and run the application on the system\u2019s hardware model deployed on an FPGA.<\/li> <\/ol> <p><strong>Required knowledge and skills:<\/strong><\/p> <ul class='wp-block-list'> <li>RTL design and FPGA implementation in SystemVerilog<\/li> <li>Good understanding of memory architectures and microcontrollers<\/li> <li>Good analytical skills<\/li> <li>Good background in computer architecture<\/li> <li>Teamwork and Git<\/li> <\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul class='wp-block-list'> <li>Scientific curiosity<\/li> <li>Good communication skills<\/li> <li>Advanced English&nbsp;<\/li> <\/ul> <p><strong>Type of work:<\/strong> 40% theory analysis, 60% SW\/HW co-design and simulation<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Michele Caon, Prof. David Atienza<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch; michele.caon@epfl.ch; david.atienza@epfl.ch?subject=Implementation of an Accelerator based on Near-Memory Computing IPs for a RISC-V-based Microcontroller'>davide.schiavone@epfl.ch; michele.caon@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project557minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Michele Caon, Prof. David Atienza<br>\";<\/script>\n<span id=project557><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Michele Caon, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project557',project557); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor548><\/a><b><span style='font-size: 20px;'>X-HEEP Accelerators: design, verification, and integration of a general purpose co-processor\/accelerator based on RISC-V for edge-computing SoCs<\/b><br><script>var project548=\"Microcontrollers (MCUs) are used in a wide range of applications ranging from sensors monitoring all the way to robotics and automotive. Despite typically lower in performance, they are usually preferred over custom circuits and FPGAs thanks to their versatility and easy programmability via software routines typically written in the C language. <br \/><br \/>Thanks to their versatility, MCUs are typically chosen as edge computing platform. However, deploying edge computing kernels (e.g. signal processing, neural networks, etc.) on resource-constrained and power-limited devices poses serious challenges in delivering real-time performance. In addition, edge computing platforms are battery powered, thus achieving high energy efficiency is needed to increase the battery life-time. &nbsp;<br \/><br \/>For these reasons, edge computing platforms are typically extended with accelerators. However, accelerators are typically built around a specific kernel to maximize performance and energy efficiency, and thus not versatile, which makes its design and verification more expensive and not versatile.<br \/><br \/>The goal of this thesis is to design and implement a general purpose accelerator for edge-computing applications based on RISC-V. The accelerator can be either based on coarse-grain reconfigurable arrays (CGRAs), or on Graphic Processing Units (GPUs) and it will work together a RISC-V core. <br \/><br \/>The accelerator will exploit the data, instruction or thread parallelism to significantly increase the performance and the energy efficiency of typical edge-computing benchmarks. <br \/><br \/>The accelerator will be integrated in X-HEEP, (eXtendable Heterogeneous Energy-Efficient Platform), an open-source, configurable, and extensible single-core RISC-V MCU, sponsored by the EcoCloud Sustainable Computing center of EPFL, and&nbsp; developed at the Embedded Systems Laboratory (ESL) jointly with other EPFL laboratories.<br \/><br \/> The expected outcomes of this thesis are:  &nbsp; <ul>  \t<li>Design\/Extend the accelerator to improve the performance\/area\/power figures or the programmability of the accelerator to make it more general-purpose<\/li>  \t<li>Extend the open-source RISC-V X-HEEP with an accelerator<\/li>  \t<li>Benchmark the performance of the accelerator of kernels on the FPGA<\/li>  \t<li>Compare the performance of the accelerator against the X-HEEP CPU on a given set of kernels on the FPGA<\/li>  \t<li>Provide performance\/power\/area (PPA) figures of the accelerator in tsmc65 LP technology<\/li>  \t<li>Compare the accelerator with other general-puropose or fixed-function state-of-the-art accelerators<\/li> <\/ul> Throughout the project, the student will learn: <ul>  \t<li>How to design or extend an accelerator that can perform general-purpose kernels oriented to signal-processing<\/li>  \t<li>How to extend the RISC-V X-HEEP platform with the designed accelerator, this imposes constraints on the accelerator HW and SW interface<\/li>  \t<li>How to analyze the PPA figures of the accelerator in a given technology throught the ASIC flow<\/li>  \t<li>How to compare the proposed solution against state-of-the-art accelerators<\/li>  \t<li>How to work with git repositories and in a team of people all contributing to the same project.<\/li> <\/ul>The project will be carried out at the ESL at EPFL, one of the world's top-class universities including EcoCloud&rsquo;s technical support. ESL is an active group (24 Ph.D. students among 45 members) involved in many research aspects. The student will be under the supervision of Prof. David Atienza and Dr. Davide Schiavone.<br \/><br \/><strong>Project objectives:<\/strong> <ol>  \t<li>Understanding the X-HEEP microcontroller, how it works, and learning how IPs are integrated. Understand how the configuration script of the bus and CPU is implemented.<\/li>  \t<li>Understanding how to design accelerators in SystemVerilog and how to extend the MCU with that accelerator.<\/li>  \t<li>Understanding the ASIC flow to analyze the PPA figures.<\/li>  \t<li>Validation of the proposed accelerator with C tests.<\/li>  \t<li>Comparison against state-of-the-art solutions over a set of applications.<\/li> <\/ol> <strong> Required knowledge and skills:<\/strong> <ul>  \t<li>RTL design and in any HDL (SystemVerilog is preferred and is going to be used throughout the project)<\/li>  \t<li>Python basic skills and algorithm implementations<\/li>  \t<li>Good understanding of memory architectures and microcontrollers<\/li>  \t<li>Good analytical skills<\/li>  \t<li>Good background in computer architecture and algorithms<\/li>  \t<li>Teamwork and git<\/li> <\/ul> <strong>&nbsp;<\/strong>  <strong>Appreciated skills:<\/strong> <ul>  \t<li>Scientific curiosity<\/li>  \t<li>Good communication skills<\/li>  \t<li>Advanced English<\/li> <\/ul> <br \/> <strong>Type of work:<\/strong> 20% theory analysis, 80% design and simulation<br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Prof. David Atienza<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch;david.atienza@epfl.ch?subject=X-HEEP Accelerators: design, verification, and integration of a general purpose co-processor\/accelerator based on RISC-V for edge-computing SoCs'>davide.schiavone@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project548minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Prof. David Atienza<br>\";<\/script>\n<span id=project548><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project548',project548); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor540><\/a><b><span style='font-size: 20px;'>Towards intelligent laser welding: automatic control of processing quality using reinforcement learning and deep neural architectures<\/b><br><script>var project540=\"Laser welding is a critical technology for many key economic sectors, including aerospace, automotive, medical industries. The on-growing importance of this technology requires insurance of processing quality, which is lacking in modern industrial welding systems. The reason is a non-linear nature of light - mater interactions, which complicates the development of quality control with robust operation under varying manufacturing conditions and material properties. Recent advances in artificial intelligence (AI) and machine learning (ML) are a potential solution that allows recovering complex data regularities in a self-learning manner. The existing challenges in AI\/ML control developments are the self-learning\/adaption mechanisms and search for optimal solutions under uncertainty constraints. The development of such AI\/ML control is the prime goal of this project.<br \/>Keywords: Laser welding, reinforcement learning, control, self-learning, quality in laser processing<br \/><br \/><strong>Workplan<\/strong><br \/><br \/>As a part of ML\/AI team, you will work in the interdisciplinary field of laser physics, laser technology, sensors and AI\/ML. Your main activity will be devoted to development of AI\/ML control algorithms (80% of time), while the efficiency of those will be supported by experiments, involving unique laser welding equipment (20% of time). During your work, you will touch such topics, as machine learning, probability theory, topology and non-linear dynamics. You will push your algorithm towards a complete autonomous learning of the laser welding. In particular you will develop unsupervised self-learning procedures, that will bring the algorithm through the first baby-steps to a complete mastering of laser welding within a limited time without any himan interventions.<br \/><br \/><strong>Required skills<\/strong><br \/><br \/>We are looking for the internships and master student with the background in electrical and electronic engineering, applied mathematics, computational science or engineering. The position assumes the basic knowledge of control and probability theories (in latter, in particular, the concepts of reinforcement leanring are benefitial). The hands on and awareness of the main concepts of machine learning is preferable as well. The position assumes programming skills in python and\/or C++. Experience with real-time systems is benefitial. <br \/><br \/><strong>Languages:<\/strong> English (Advanced)<br \/><br \/><strong>Location: <\/strong>Empa Thun<br \/><br \/><strong>Project\/contact:<\/strong> This work is a part of intensive research in intelligent industrial automation that aims to develop digital twins for laser material processing. More details about current activities can be found on the group webpage:<br \/> <br \/><a href='https:\/\/www.empa.ch\/web\/s204'>https:\/\/www.empa.ch\/web\/s204<\/a><br \/><p>Further technical details about the project and applications can be sent\/discussed with <a href='mailto:sergey.shevchik@empa.ch'>Dr. Sergey Shevchik <\/a><\/p><p>&nbsp;<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Sergey Shevchik, Prof. David Atienza<br> Contact email: <a href='mailto:sergey.shevchik@empa.ch; david.atienza@epfl.ch?subject=Towards intelligent laser welding: automatic control of processing quality using reinforcement learning and deep neural architectures'>sergey.shevchik@empa.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project540minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Sergey Shevchik, Prof. David Atienza<br>\";<\/script>\n<span id=project540><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Sergey Shevchik, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project540',project540); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=http:\/\/www.empa.ch\/ target=_blank title='EMPA (Eidgen\u00f6ssische Materialpr\u00fcfungs- und Forschungsanstalt)'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/29.png width=70 alt='EMPA (Eidgen\u00f6ssische Materialpr\u00fcfungs- und Forschungsanstalt)'><\/a><\/td><\/tr><tr><td colspan=2><h3>Master or Semester Projects<br><br><\/h3><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor707><\/a><b><span style='font-size: 20px;'>Custom Hardware Design of Posit Arithmetic Operators for Low-Precision Computing<\/b><br><script>var project707=\"<p>Computer arithmetic is the foundation of all numerical computation in digital systems. It encompasses the algorithms and hardware structures used to perform basic operations such as addition, multiplication, and division on binary-encoded numbers. These operations are implemented using digital logic circuits that manipulate bits through combinational and sequential logic. The choice of number representation, such as fixed-point or floating-point, directly impacts the precision, performance, and energy efficiency of arithmetic units. As applications increasingly demand low-power and high-throughput computation, especially in embedded and signal processing systems, the design of efficient arithmetic hardware has become a critical area of research and innovation.<\/p> <p>Posit arithmetic is an emerging number representation system designed to enhance numerical accuracy and efficiency, particularly in low-precision computing applications. Unlike traditional floating-point formats, posits offer a tapered precision scheme that allocates more bits to the exponent or fraction depending on the magnitude of the number. This dynamic allocation enables better precision near unity and a wider dynamic range, making posits especially attractive for signal processing and embedded AI applications.<\/p> <p>This project aims to develop a library of parameterized arithmetic modules for core posit arithmetic operations in SystemVerilog. These modules will enable flexible and efficient hardware design for signal processing or machine learning tasks. The designs will emphasize modularity, configurability (e.g., posit size and exponent size), and synthesis efficiency.<\/p> <p>The project will be carried out at the ESL at EPFL, one of the world's top-class universities. ESL is an active group (24 Ph.D. students among 45 members) involved in many research aspects. The student will be under the supervision of Mr. Tommaso Terzano, Dr. David Mallas&eacute;n Quintana and Prof. David Atienza. <\/p> <p><strong>Project objectives:<\/strong> <\/p> <p>The objectives can be adapted to both a semester project or master thesis. Depending on the workload, some will be mandatory and others optional. The full list is the following: <\/p> <ol>     <li>         <p>Understanding posit arithmetic and studying <a rel='noopener noreferrer nofollow' href='https:\/\/github.com\/RaulMurillo\/Flo-Posit' target='_blank'><span style='color: #1155cc'>previous<\/span><\/a><span style='color: #0e101a'> hardware implementations of posit units.<\/span>         <\/p><\/li><li>         <p>Designing modular and configurable implementations of posit units in SystemVerilog:         <\/p>         <ol>             <li>                 <p>Addition\/Subtraction                 <\/p>             <\/li>             <li>                 <p>Multiplication                 <\/p>             <\/li>             <li>                 <p>Multiply-Accumulate (MAC)                 <\/p>             <\/li>             <li>                 <p>Division                 <\/p>             <\/li>             <li>                 <p>Square root                 <\/p>             <\/li>             <li>                 <p>Conversion to and from integers                 <\/p>             <\/li>             <li>                 <p>Quire MAC                 <\/p>             <\/li>         <\/ol>     <\/li>     <li>         <p>Providing diagrams and documentation for each hardware block.         <\/p>     <\/li>     <li>         <p>Verifying the functionality of the posit units, for <a rel='noopener noreferrer nofollow' href='https:\/\/github.com\/davidmallasen\/arithmetic_units' target='_blank'><span style='color: #1155cc'>example<\/span><\/a><span style='color: #0e101a'> using cocotb.<\/span>         <\/p>     <\/li>     <li>         <p>Integrating the developed units into a full arithmetic unit, for example <a rel='noopener noreferrer nofollow' href='https:\/\/github.com\/openhwgroup\/cvfpu' target='_blank'><span style='color: #1155cc'>CVFPU<\/span><\/a><span style='color: #0e101a'>.<\/span>         <\/p>     <\/li>     <li>         <p>Providing ASIC synthesis results of the developed blocks and comparing them with existing implementations.         <\/p>     <\/li>     <li>         <p>Implement pipelined versions of selected operators and analyze trade-offs.         <\/p>     <\/li> <\/ol> <p>Note: Evaluation will be based on the overall performance and dedication of the student. If the student successfully completes the objectives, the next step would be participating in the research activity, which will be based on the student&rsquo;s interests. <\/p> <p><strong>Required knowledge and skills:<\/strong> <\/p> <ul>     <li>         <p>Excellent RTL design skills, ideally in SystemVerilog, demonstrated through past projects.         <\/p>     <\/li>     <li>         <p>Previous experience or knowledge of computer arithmetic is a plus.         <\/p>     <\/li>     <li>         <p>Basic Python programming.         <\/p>     <\/li>     <li>         <p>Ability to work consistently, independently, and communicate effectively in English.         <\/p>     <\/li>     <li>         <p>Familiarity with Git version control.         <\/p>     <\/li> <\/ul> <p>     <br \/><strong>Type of work:<\/strong> 20% theory analysis, 60% design and simulation, 20% verification and documentation <\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Mr. Tommaso Terzano, Dr. David Mallas\u00e9n Quintana and Prof. David Atienza<br> Contact email: <a href='mailto:tommaso.terzano@epfl.ch;david.mallasen@epfl.ch;david.atienza@epfl.ch?subject=Custom Hardware Design of Posit Arithmetic Operators for Low-Precision Computing'>tommaso.terzano@epfl.ch;david.mallasen@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project707minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Mr. Tommaso Terzano, Dr. David Mallas\u00e9n Quintana and Prof. David Atienza<br>\";<\/script>\n<span id=project707><b>Lab: <\/b>ESL<br><b>Sections: <\/b>STI<br><b>Supervisor<\/b>: Mr. Tommaso Terzano, Dr. David Mallas\u00e9n Quintana and Prof. David Atienza<br> <a href=#_ onclick=opendesc('project707',project707); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor630><\/a><b><span style='font-size: 20px;'>Enhancing the Efficiency and Accuracy of Radio Interferometry Kernels using CGRA-ME Framework<\/b><br><script>var project630=\"<div><strong><br \/><\/strong><\/div><strong>Background<\/strong><div><strong><br \/><\/strong>The Square Kilometer Array (SKA) is one of the most ambitious international scientific projects ever undertaken. It aims to build the world's largest radio telescope, with sites in Australia and South Africa. The SKA will enable unprecedented observations of the universe across a wide range of radio frequencies, facilitating groundbreaking research in fields such as astrophysics, cosmology, and fundamental physics. The sheer scale of the SKA, which will generate an enormous amount of data (up to 10 Tb\/s at the input), presents significant computational challenges, particularly in processing and analyzing the data streams in real time. To address these challenges, the SEAMS (sustainable &amp; energy aware methods for SKA observatory) project has been initiated. SEAMS is focused on developing energy-efficient and scalable architectures for processing the massive data generated by SKA. The project seeks to optimize key computational kernels used in radio interferometry, such as Fast Fourier Transform (FFT), convolution, and deconvolution, which are essential for converting raw telescope data into meaningful data products. High-Performance Computing (HPC) is critical to the SKA's success, and optimizing the performance and energy efficiency of these kernels is a primary goal.<p><strong>Project Motivation<\/strong><\/p><p>One promising approach to achieving these objectives is the use of Coarse-Grained Reconfigurable Architectures (CGRA), which offer a balance between the flexibility of general-purpose processors and the efficiency of custom hardware accelerators. The CGRA-ME framework is an open-source tool that allows for the design, configuration, and evaluation of CGRA architectures. By leveraging CGRA-ME, this project aims to enhance the performance and energy efficiency of HPC kernels critical to the SKA and SEAMS projects.<\/p><p><strong>Project Objectives<\/strong><\/p><p>1. Configure CGRA-ME and Evaluate the Performance and Efficiency of a Commonly Used HPC Kernel in Radio Interferometry<br \/>Objective Details:<br \/>o Implement and validate the execution of a basic kernel such as FFT, convolution, and\/or deconvolution using CGRA-ME framework.<br \/>o Evaluate the performance and energy efficiency of these kernels will be evaluated on a basic CGRA architecture to establish a baseline for further optimization.<br \/>o Compare performance results against other platforms.<br \/>o Identify bottlenecks in the computation and propose improvements.<\/p><p>2. Design a New Processing Element for Variable Precision Arithmetic Operators and Integrate it into the Basic CGRA Architecture<br \/>Objective Details:<br \/>o Radio interferometry computations often require varying levels of precision depending on the specific operation and data characteristics. Design a new processing element (PE) within the CGRA architecture that supports variable precision arithmetic. Explore different precision configurations.<br \/>o Test the correct functioning of the design and quantify the precision drop\/increment and performance.<br \/>o Explore other CGRA architecture aspects, such as communication, private memories, etc.<\/p><p>3. Evaluate the Precision, Performance, and Energy Trade-offs of the Design<br \/>Objective Details:<br \/>o The final objective of the project is to thoroughly evaluate the trade-offs between precision, performance, and energy consumption for the newly designed processing element.<br \/>o This evaluation will involve comparing the variable precision PE's performance with that of fixed-precision PEs in executing the chosen HPC kernels.<\/p><p><strong>Required Knowledge and Skills<\/strong><\/p><p>&bull; Hardware Design: Experience with hardware description languages (HDLs) such as Verilog or VHDL.<br \/>&bull; Reconfigurable Architectures: Understanding of reconfigurable architectures such as FPGAs and CGRAs.<br \/>&bull; Analysis: Knowledge of techniques to analyze and optimize energy consumption in hardware designs.<\/p><p><strong>Type of Work<\/strong><\/p><p>&bull; Theoretical Analysis (30%): This will involve the design of the new processing element and the theoretical exploration of precision scaling in arithmetic operations.<br \/>&bull; Design and Implementation (50%): Hands-on configuration and modification of the CGRA-ME framework, integration of the new processing element, and execution of kernel evaluations.<br \/>&bull; Testing and Evaluation (20%): Extensive testing, validation, and analysis of performance, precision, and energy trade-offs.<br \/>Expected Outcomes<br \/>&bull; A new processing element for variable precision arithmetic, integrated into the CGRA, validated, and tested.<br \/>&bull; Comprehensive evaluation results showing the trade-offs between precision, performance, and energy efficiency, offering valuable insights for the SKA and SEAMS projects.<\/p><p><strong>Learning objectives<\/strong><\/p><p>&bull; Research Skill: The student will develop the ability to synthesize information from various sources, apply theoretical knowledge to practical problems, and innovate in their design approach.<br \/>&bull; Technical Skills: The student will gain hands-on experience with Coarse-Grained Reconfigurable Architectures (CGRA), particularly using the CGRA-ME framework, including configuring, modifying, and optimizing these systems for specific applications.<br \/>&bull; Analytical Skills: The student will develop the ability to critically evaluate and analyze the trade-offs involved in different design choices, particularly regarding precision, performance, and energy efficiency.<br \/>&bull; Communication Skills: The student will learn how to present these trade-offs clearly, which is essential for making informed design decisions in engineering projects.<br \/>&bull; Problem-Solving Skills: Throughout the project, the student will engage in problem-solving and research, learning how to approach complex engineering challenges methodically.<\/p><p>This project will be conducted under the supervision of experts in hardware design, energy efficient HPC, and optimization from the Embedded Systems Laboratory (ESL), providing the student with an opportunity to contribute to cutting-edge research in energy-efficient computing: Dr. Denisa-Andreea Constantinescu denisa.constantinescu@epfl.ch, Rub&eacute;n Rodr&iacute;guez &Aacute;lvarez ruben.rodriguezalvarez@epfl.ch, Dr. Giovanni Ansaloni giovanni.ansaloni@epfl.ch, and Prof. David Atienza Alonso david.atienza@epfl.ch<br \/> <br \/><strong>References<\/strong><\/p><p>1. SEAMS Project: https:\/\/seams-project.com\/<br \/>2. CGRA-ME Framework: https:\/\/cgra-me.ece.utoronto.ca\/<\/p><\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>ESL<br><b>Supervisor<\/b>: Dr. Denisa-Andreea Constantinescu; Rub\u00e9n Rodr\u00edguez \u00c1lvarez; Dr. Giovanni Ansaloni; Prof. David Atienza Alonso <br> \";<\/script>\n<script>var project630minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>ESL<br><b>Supervisor<\/b>: Dr. Denisa-Andreea Constantinescu; Rub\u00e9n Rodr\u00edguez \u00c1lvarez; Dr. Giovanni Ansaloni; Prof. David Atienza Alonso <br>\";<\/script>\n<span id=project630><b>Lab: <\/b>ESL<br><b>Sections: <\/b>ESL<br><b>Supervisor<\/b>: Dr. Denisa-Andreea Constantinescu; Rub\u00e9n Rodr\u00edguez \u00c1lvarez; Dr. Giovanni Ansaloni; Prof. David Atienza Alonso <br> <a href=#_ onclick=opendesc('project630',project630); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor620><\/a><b><span style='font-size: 20px;'>Integration and optimization of Ultra Low Power CPU on the X-HEEP platform targeting implantable devices\r\n<\/b><br><script>var project620=\"<h3>Project Description<\/h3> <p style='text-align: justify'>The Embedded Systems Laboratory (ESL) of EPFL has been at the forefront of developing open-source, energy-efficient computing platforms, most notably the eXtendable Heterogeneous Energy-Efficient Platform (<a href='https:\/\/github.com\/esl-epfl\/x-heep' target='_blank'>     <strong>X-HEEP<\/strong>   <\/a>). X-HEEP is a versatile, RISC-V-based microcontroller designed to target both small-scale and high-performance applications. This platform provides a customizable and extendable MCU, allowing users to integrate their own accelerators without modifying the core microcontroller architecture. X-HEEP has been proven effective in various implementations, from FPGA to ASIC, showcasing its adaptability and performance efficiency in diverse scenarios.&nbsp;<\/p> <p style='text-align: justify'>Although X-HEEP&rsquo;s energy efficiency has been proven to match the requirements of wearable devices, in the domain of implantable technology, there is an increasing demand for devices that are not only energy efficient, but also extremely low power. Wireless power delivery often found in implantable devices needs to operate at a constrained power budget to avoid damage to tissue due to overheating. This forces devices to go down from hundreds of mW (typical on wearable devices) to tens of &micro;W:   <strong>a four-orders-of-magnitude drop!<\/strong> <\/p> <p style='text-align: justify'>In this light, the X-HEEP project is opening a branch to attain unprecedented power efficiency while retaining the configurability, extendibility, and re-programmability that make it a desirable processing platform. The proposed project will kick off this path by instantiating an extremely low-power CPU on X-HEEP: the   <a href='https:\/\/github.com\/olofk\/serv' target='_blank'>SERV processor<\/a>. Then, the integration process will be optimized to the requirements of biosignal processing, usually characterized by time sparsity and medium amplitude resolution.&nbsp;<\/p> <p style='text-align: justify'>The project can be adapted to either a Master-level Semester Project or a Master&rsquo;s Thesis.&nbsp;<\/p> <p style='text-align: justify'>It will be carried out at the ESL at EPFL. ESL is an active group (22 Ph.D. students among 45 members) involved in many research aspects. The student will be under the supervision of Mr. Juan Sapriza, Dr. Davide Schiavone, and Prof. David Atienza.<\/p> <h3 style='text-align: justify'>Project Objectives<\/h3> <ol>   <li>Integrate the SERV processor into the X-HEEP platform.&nbsp;<\/li>   <li>Perform RTL-level adaptations to improve its performance in the biosignal processing context (for example, implementing optimized low-bit-width data computing as an addition of 4-bit data).&nbsp;<\/li>   <li>Characterize the improvements and validate the full functionality of the CPU.&nbsp;<\/li>   <li>(Optional) Help integrate the CPU as the main processor of the latest X-HEEP chip tape-out.&nbsp;<\/li> <\/ol> <h3 style='text-align: justify'>Required knowledge and skills:<\/h3> <ul>   <li>Advanced knowledge on computer architecture and RTL<\/li>   <li>Creativity, autonomy and scientific rigor<\/li>   <li>Embedded C and\/or Assembly<\/li><li>Confidence working with Linux systems<\/li>   <li>Git<\/li> <\/ul> <h3 style='text-align: justify'>Type of work: <\/h3> <p style='text-align: justify'>60% Implementation, 40% Characterization and testing<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Juan Sapriza, Dr. Davide Schiavone, Prof. David Atienza<br> \";<\/script>\n<script>var project620minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Juan Sapriza, Dr. Davide Schiavone, Prof. David Atienza<br>\";<\/script>\n<span id=project620><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Juan Sapriza, Dr. Davide Schiavone, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project620',project620); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><tr><td colspan=2><h3>Semester Projects<br><br><\/h3><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor770><\/a><b><span style='font-size: 20px;'>FPGA implementation and validation of a Compute-Memory architecture for DNN inference applications <\/b><br><script>var project770=\"<p>The objective of this project is to create a demonstrator of Faraday,  a configurable SRAM-based compute-near-memory architecture developed at  the Embedded Systems Laboratory of EPFL. Targeting the Xilinx  UltraScale+ ZCU104 FPGA board, the student will synthesize and deploy  the Compute Memory architecture integrated with a 32-bit RISC-V platform  and off-chip DRAM. <\/p><p>The student will first familiarize themselves  with the Faraday architecture, its memory organization, and the  software framework used to generate C applications for the system. They  will then study the application generation flow and the mechanisms used  to store the parameters of DNN models in the external DRAM and transfer  them to the Compute Memory banks for execution. <\/p><p>The project main  objective will be to validate the correct execution of different DNN  inference workloads on the FPGA board, from the input data to the final  inference result. The student will progressively test the different  components of the execution flow and verify the FPGA results against a  software reference, with the goal of demonstrating a complete end-to-end  DNN inference running on Faraday. <\/p><p>This implementation will also  allow the student to perform runtime exploration of the DNN models over  different Compute Memory configurations. By varying the number and  organization of the Compute Memory banks, the student will investigate  the impact of the architecture configuration on inference execution time  and assess the potential of Faraday for real-time DNN inference  applications. <\/p><p><strong>Tasks description <\/strong><\/p><p>Required tasks (to be completed for a passing grade of 4.00\/6.00) <\/p><p>1.  Study the organization of the Compute Memory architecture, its memory  banks, data representation, execution flow, and interaction with the  32-bit RISC-V platform. <\/p><p>2. Build and synthesize a Faraday  architectural instance integrated with the RISC-V processor and off-chip  DRAM for the Xilinx UltraScale+ ZCU104 development board. <\/p><p>3.  Study the software framework used to generate the C applications  executed on the RISC-V platform and understand how DNN models are mapped  to the Faraday architecture. <\/p><p>4. Execute a CNN model (AlexNet) on  the ZCU104 board and verify the correctness of the inference results by  comparing them with a software reference implementation.  <\/p><p>The accomplishment of each of the following will add a further point to the final grade:  <\/p><p>5.  Evaluate the execution of the selected DNN models using different  Compute Memory configurations and analyse their impact on inference  latency and overall system performance.  <\/p><p>6. Hide the latency of the external DRAM accesses on the execution time of inference on Faraday.  <\/p><p>7.  Interface a camera with the Xilinx UltraScale+ ZCU104 and implement the  required data path to acquire images, store them in the external DRAM,  and use them as inputs to a DNN running on Faraday. Demonstrate an  end-to-end image recognition application on the FPGA. (optional) <\/p><p><strong>Required knowledge and skills <\/strong><\/p><p>&middot; Good HDL programming skills (VHDL and Verilog). <\/p><p>&middot; Good programming skills in C and Python. <\/p><p>&middot; Basic knowledge of computer architecture and embedded systems. <\/p><p>&middot; Scientific curiosity. <\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Gr\u00e9goire Eggermann, Dr. Giovanni Ansaloni and Prof. David Atienza<br> Contact email: <a href='mailto:gregoire.eggermann@epfl.ch;giovanni.ansaloni@epfl.ch;david.atienza@epfl.ch?subject=FPGA implementation and validation of a Compute-Memory architecture for DNN inference applications '>gregoire.eggermann@epfl.ch;giovanni.ansaloni@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project770minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Gr\u00e9goire Eggermann, Dr. Giovanni Ansaloni and Prof. David Atienza<br>\";<\/script>\n<span id=project770><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Gr\u00e9goire Eggermann, Dr. Giovanni Ansaloni and Prof. David Atienza<br> <a href=#_ onclick=opendesc('project770',project770); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor765><\/a><b><span style='font-size: 20px;'>Automatic geolocation of acoustic events using the VersaSens platform<\/b><br><script>var project765=\" <p>The objective of this project is to develop and experimentally validate a system for the automatic detection and localization of acoustic events in indoor environments using the VersaSens platform.<\/p> <p>VersaSens, developed at the Embedded Systems Laboratory (ESL), is a modular, multimodal, extendable, and reconfigurable Edge AI platform. Its architecture enables the integration of add-on sensing modules together with embedded sensing and processing capabilities, providing a flexible platform for the development of distributed acoustic sensing applications.<\/p> <p>The project will investigate the use of multiple synchronized VersaSens devices equipped with stereo acoustic sensors to estimate the position of an acoustic source within a known indoor environment. Localization will exploit complementary acoustic information, including time of arrival\/time difference of arrival, source orientation, and signal intensity.<\/p> <p>A central objective is to characterize the <strong>accuracy, robustness, and resilience of acoustic event localization<\/strong> and to investigate the number and spatial configuration of sensing devices required to achieve reliable localization for a given room geometry. Particular attention will be given to inter-device synchronization, sensor placement, environmental conditions, and the amount of information that must be transmitted from the embedded devices to a host computer.<\/p> <p>This project contributes to an ongoing ESL research effort targeting the detection and localization of cough events using the VersaSens platform. It will build upon previous developments within the laboratory, including firmware for the synchronization of multiple VersaSens devices, sound-localization algorithms based on orientation, and embedded cough-detection algorithms.<\/p> <p>The longer-term objective of this research is to support the detection and localization of cough events potentially associated with tuberculosis in clinical waiting rooms, with the aim of supporting patient triage.<\/p> <p><strong>Mandatory tasks<\/strong><\/p> <p>Completion of these tasks is required to pass the exam and obtain a grade of 4.<\/p> <ol class='wp-block-list'><li>Conduct a literature review of acoustic source-localization techniques and become familiar with the VersaSens platform and microphone characteristics.<\/li><li>Develop the VersaSens firmware to control and synchronize at least 3 VersaSens devices with acoustic sensors in stereo mode. <strong>Characterize the synchronization accuracy and timing variability between devices.<\/strong><\/li><li>Develop a computer interface providing a visual representation of the room from given dimensions, including the positions and orientations of the VersaSens devices.<\/li><li>Implement Bluetooth communication between the synchronized VersaSens devices and the computer interface to collect the acoustic information required for localization.<\/li><li>Develop the signal-processing and coregistration pipeline to estimate and visualize the sound-source position in real time, combining <strong>time-of-arrival\/time-difference-of-arrival, orientation, and signal-intensity information<\/strong>.<\/li><li>Optimize the settings, including number and location of VersaSens devices and orientation of the microphones, given a room geometry and\/or the region of interest in the room.<\/li><li>Develop preliminary on-device signal processing to transmit only the information necessary for acoustic-event localization and reduce communication requirements.<\/li><li>Conduct a systematic validation experiment simulating the real-life application using known source positions. <strong>Quantify localization accuracy and evaluate the influence of sensor number\/configuration and environmental conditions.<\/strong><\/li><li>Deliver a complete documentation package for the VersaSens GitHub repository, including datasets, software, firmware, experimental procedures, and results.<\/li><\/ol> <p><strong>Optional tasks<\/strong><\/p> <p>Completion of each task will contribute additional points to the final grade (up to 6 total).<\/p> <ol class='wp-block-list'><li>Extend the interface to support <strong>complex room geometries<\/strong> and adapt the localization approach to consider walls and corners.<\/li><li>Conduct a validation experiment simulating the real-life application in complex rooms and quantify the resulting geolocation accuracy.<\/li><li>Discuss and validate the case of oriented sound events and\/or oriented areas of interested. For example, people cough towards a direction and the orientation of chairs in a room may be known.<\/li><\/ol> <p><strong>Type of work<\/strong><\/p> <ul class='wp-block-list'><li>30% firmware development<\/li><li>30% software development and signal-processing development<\/li><li>20% Experimental acquisition in different conditions to mimic real-life application.<\/li><li>10% Performance evaluation through systematic testing of the platform.<\/li><li>10% Preparation and delivery of the complete documentation package.<\/li><\/ul> <p><strong>Desired skills:<\/strong><\/p> <ul class='wp-block-list'><li>Strong analytical skills<\/li><li>Strong background in signal processing<\/li><li>Experience in embedded systems development<\/li><li>Teamwork and git<\/li><\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul class='wp-block-list'><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English<\/li><\/ul> <br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Doctoral candidate Wensi Zhang, Dr. Taraneh Aminosharieh Najafi, Dr. J\u00e9r\u00f4me Thevenot<br> Contact email: <a href='mailto:jerome.thevenot@epfl.ch; taraneh.aminoshariehnajafi@epfl.ch; wensi.zhang@epfl.ch?subject=Automatic geolocation of acoustic events using the VersaSens platform'>jerome.thevenot@epfl.ch; taraneh.aminoshariehnajafi@epfl.ch; wensi.zhang@epfl.ch<\/a><br>\";<\/script>\n<script>var project765minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Doctoral candidate Wensi Zhang, Dr. Taraneh Aminosharieh Najafi, Dr. J\u00e9r\u00f4me Thevenot<br>\";<\/script>\n<span id=project765><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Doctoral candidate Wensi Zhang, Dr. Taraneh Aminosharieh Najafi, Dr. J\u00e9r\u00f4me Thevenot<br> <a href=#_ onclick=opendesc('project765',project765); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor762><\/a><b><span style='font-size: 20px;'>Scalable Memory-Mapped Systolic Array Cluster for LLM Inference<\/b><br><script>var project762=\"<p>This project investigates how systolic-arrays can be integrated as modular, memory-mapped accelerator tile that can be replicated to scale up the computing capabilities of a multi-accelerator system. Starting from the existing TiC-SAT architecture, the project will explore a loosely coupled organization in which multiple systolic-array instances can be independently controlled and operate concurrently.<\/p> <p>A systolic array consists of a network of tightly coupled data-processing units, referred to as Processing Elements (PEs), through which data flows in a structured manner. Each PE performs a small portion of the overall computation\u2014typically multiply-accumulate (MAC) operations\u2014and forwards intermediate results to neighboring elements. This architecture enables massive parallelism and high data reuse. By significantly reducing accesses to off-chip memory, systolic arrays achieve exceptional energy efficiency for matrix-intensive workloads, which form the computational backbone of deep learning models and large language models (LLMs). However, as the computational demands of these workloads increase, the performance achievable with a single accelerator becomes increasingly limited.<\/p> <p>To enable scale-up through accelerator replication, the project will integrate TiC-SAT into a memory-mapped architecture based on standard AXI interfaces. The resulting accelerator will be evaluated on a Xilinx Zynq UltraScale+ FPGA platform, where multiple independently controlled systolic-array instances will be deployed to demonstrate concurrent execution and assess the scalability of the proposed architecture. The project thus provides a first step toward scalable many-accelerator architectures for high-performance LLM inference.<\/p> <p><strong>Tasks description<\/strong><\/p> <ol> <li>Understand TiC-SAT.<\/li> <\/ol> <p>Study the TiC-SAT systolic array RTL implementation, including the invocation mechanism, accelerator interfaces, dataflow, control logic, and available test workloads.<\/p> <ol start='2'> <li><strong>Memory-mapped systolic-array accelerator.<\/strong><\/li> <\/ol> <p>Integrating TiC-SAT as a memory-mapped accelerator. Design an AXI-compatible hardware\/software interface providing accelerator configuration, execution control, and status monitoring. Define the required slave registers and develop control logic, based on a finite-state machine, to manage the internal systolic-array operations. Develop an RTL testbench to verify the memory-mapped interface and accelerator functionality against the existing functional model.<\/p> <ol start='3'> <li><strong>FPGA emulation framework.<\/strong><\/li> <\/ol> <p>Integrate the memory-mapped accelerator into the Programmable Logic of a Xilinx Zynq UltraScale+ platform and develop a complete PS\/PL hardware\/software co-design framework. Implement the software support required to configure, launch, synchronize, and validate accelerator execution from a Linux-based operating system running on the Processing System. Characterize the accelerator in terms of functional correctness, execution time, and FPGA resource utilization.<\/p> <ol start='4'> <li>Multi-systolic-array architecture.<\/li> <\/ol> <p>Extend the FPGA architecture with at least two independently addressable and controllable systolic-array instances. Develop the required AXI interconnect and software infrastructure and demonstrate concurrent execution of matrix operations on multiple accelerators. Compare the execution of representative workloads using a single SA and multiple SAs and analyze the scalability of the architecture, including the impact of shared memory and interconnect resources.<\/p> <p><strong>Project objectives<\/strong><\/p> <p>The fulfillment of the mandatory objectives will result in a passing grade of 4.0. The completion of the two optional objectives will increase the grade by 1.0 points each.<\/p> <ul> <li>[Mandatory] Extend the existing RTL systolic-array description with an AXI memory-mapped control interface, allowing software to configure, launch, and monitor accelerator operations. Define the required control and status registers and implement the corresponding accelerator control logic.<\/li> <li>[Mandatory] Develop a testbench for validating the memory-mapped accelerator interface in RTL and verifying accelerator results against the existing functional model.<\/li> <li>[Mandatory] Develop a complete FPGA emulation environment on a Xilinx Zynq UltraScale+ platform capable of executing representative matrix operations on the systolic array from a Linux-based operating system running on the Processing System.<\/li> <li>[Mandatory] Integrate at least two independently addressable systolic-array instances into a common AXI-based memory and interconnect subsystem and demonstrate their concurrent operation.<\/li> <li>[Optional] Develop and evaluate software mechanisms for partitioning representative LLM matrix operations among multiple systolic-array instances and investigate different workload-distribution policies.<\/li> <li>[Optional] Characterize the performance scaling and identify system bottlenecks such as memory bandwidth, interconnect contention, accelerator utilization, and workload imbalance.<\/li> <\/ul> <p><strong>Required knowledge and skills<\/strong><\/p> <ul> <li>Proficiency in RTL design and programming (e.g., VHDL or Verilog).<\/li> <li>Experience with FPGA design and implementation.<\/li> <li>Basic knowledge of C\/C++ programming.<\/li> <li>Basic understanding of computer architecture and hardware\/software interfaces.<\/li> <li>Familiarity with embedded systems and AXI-based systems is beneficial.<\/li> <li>Strong analytical thinking and scientific curiosity.<\/li> <\/ul> <p><strong>References<\/strong><\/p> <p>[1] A. Amirshahi, J. Klein, G. Ansaloni, D. Atienza, \u201cTiC-SAT: Tightly-coupled Systolic Accelerator for Transformers\u201d,\u00a0 28th Asia and South Pacific Design Automation Conference (ASP-DAC '23), Tokyo, Japan, doi: 10.1145\/3566097.3567867 (2023)<\/p> <p>[2] J. Klein, I. Boybat, G. Ansaloni, M. Zapater, &amp; D. Atienza. \u201cWhich coupled is best coupled? an exploration of aimc tile interfaces and load balancing for CNNs\u201d, IEEE Transactions on Parallel and Distributed Systems, 35(10), 1780-1795 (2024).<\/p> <p><strong>Type of work<\/strong><\/p> <ul> <li><strong>40%<\/strong> <strong>HW design<\/strong>: Transformation of the existing tightly coupled systolic-array accelerator into a memory-mapped architecture, development of the AXI interface and control logic, RTL verification, and integration of multiple accelerator instances.<\/li> <li><strong>40% Prototyping on FPGA<\/strong>: Development of the PS\/PL hardware\/software co-design framework and deployment of single- and multi-SA architectures on a Xilinx Zynq UltraScale+ platform.<\/li> <li><strong>20% Performance Evaluation<\/strong>: Benchmarking of representative matrix operations and analysis of the scalability, FPGA resource utilization, and system-level bottlenecks of the multi-accelerator architecture.<\/li> <\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> \";<\/script>\n<script>var project762minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project762><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project762',project762); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor755><\/a><b><span style='font-size: 20px;'>Visual X-HEEP: Enhancing X-HEEP with Camera Interfaces and VGA monitors<\/b><br><script>var project755=\"Microcontrollers (MCUs) are used in a wide range of applications, ranging from sensor monitoring all the way to robotics and automotive.&nbsp; MCUs are typically chosen as edge computing platforms, processing bio-signals, audio, or images\/videos. Due to cost (area) and power constraints, the on-chip memory, typically implemented with SRAM technology, is limited (usually below 1MB), thus limiting the software and data that can be processed. That\u2019s why MCUs are often extended with off-chip memories, typically with FLASH memories. FLASH memories are non-volatile and usually connected to the MCU via serial-peripheral interfaces (SPIs).<\/p><p>X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform) is an open-source, configurable, and extensible single-core RISC-V 32b MCU, sponsored by the EcoCloud Sustainable Computing center of EPFL. It is based on many third-party open-source IPs and in-house IPs developed at the Embedded Systems Laboratory (ESL) jointly with other EPFL laboratories.<\/p><p>X-HEEP provides a framework to run applications compiled for RISC-V on a simulator (Verilator, Questasim, or VCS), on a Xilinx FPGA, and can be implemented in silicon as well. X-HEEP uses an off-chip FLASH memory connected via SPI. Accessing data to flash can be done by means of SW or HW functions to send commands to the FLASH via SPI (read or write, quad or single mode, etc.), and program the X-HEEP DMA to move data from the SPI received data to the memory (or vice versa). In addition, a simple Cache is used to reduce the Flash activities and speed up performance.<\/p><p>Due to the many options and parameters available, verification of such Flash operation is complex. For this reason, the students applying to this project will complete and verify pre-existing blocks and integrate them into X-HEEP, both in simulation and on an FPGA. <\/p><p>In addition, to further enhance the X-HEEP capabilities to interact with the external world, the student applying for this project will extend X-HEEP with a Camera Interface compatible with cameras like the OV7670, and a VGA display controller.<\/p><p>The outcome of the thesis will be published open-source in the X-HEEP repository: <a data-mce-href='https:\/\/github.com\/x-heep\/x-heep' href='https:\/\/github.com\/x-heep\/x-heep'>https:\/\/github.com\/x-heep\/x-heep<\/a><br data-mce-bogus='1'><\/p><p>The tasks will be:<\/p><ol><li aria-level='1'><b>[Verification and Integration of Flash Controller functionalities] <\/b>Complete the verification and integration of the existing Memory Mapped bridge that translates the CPU\u2019s load and stores to Flash Commands, including the interaction with the CACHE.<ol><li aria-level='2'>C functions that read\/write to different flash sectors, cross-sectors of the Flash by means of load and stores<\/li><li aria-level='2'>C functions that read\/write and write back to different flash sectors, cross-sectors using the Cache<\/li><li aria-level='2'>Integration into X-HEEP, without adding bugs, and keeping parametrizations for the final user<\/li><\/ol><\/li><li aria-level='1'><b>[Design and Integration of the Camera Interface Peripheral] <\/b>Design and verify both in Simulation (Verilator\/Questasim) and on the FPGA the Camera Interface peripheral, starting from existing examples on GitHub.&nbsp;<ol><li aria-level='2'>Writing the SystemVerilog modules that implement the Camera Interface for X-HEEP<\/li><li aria-level='2'>Verify its functionality both in simulation and on the FPGA with the OV7670 camera<\/li><\/ol><\/li><li aria-level='1'><b>[Integration of the VGA Display Peripheral] <\/b>Integrate and verify both in Simulation (Verilator\/Questasim) and on the FPGA the existing PULP VGA Controller into X-HEEP.&nbsp;<ol><li aria-level='2'>Writing the SystemVerilog modules that bridge and instantiate the VGA controller for X-HEEP<\/li><li aria-level='2'>Verify its functionality both in simulation and on the FPGA<\/li><\/ol><\/li><\/ol><p>The project will be carried out at the ESL at EPFL, one of the world's top-class universities, including EcoCloud\u2019s technical support. ESL is an active group (24 Ph.D. students among 45 members) involved in many research aspects. The student will be under the supervision of Dr. Davide Schiavone, Dr. Michele Caon, and Prof. David Atienza.<\/p><p><b>Required knowledge and skills:<\/b><\/p><ul><li aria-level='1'>RTL design in any HDL (SystemVerilog is going to be used throughout the project)<\/li><li aria-level='1'>Low-level software design (C and\/or C++ is going to be used throughout the project)<\/li><li aria-level='1'>FPGA design, synthesis, and verification (the Pynq FPGA will be used throughout the project)<\/li><li aria-level='1'>RTL simulation with QuestaSim and Verilator<\/li><li aria-level='1'>Good understanding of microcontrollers<\/li><li aria-level='1'>Good analytical skills<\/li><li aria-level='1'>Good background in computer architecture and algorithms<\/li><li aria-level='1'>Teamwork and git<\/li><\/ul><p><b>Appreciated skills:<\/b><\/p><ul><li aria-level='1'>Scientific curiosity<\/li><li aria-level='1'>Good communication skills<\/li><li aria-level='1'>Advanced English&nbsp;<\/li><\/ul><p><b>Type of work:<\/b> 10% theory analysis, 90% design and simulation&nbsp;<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Michele Caon, Prof. David Atienza<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch; michele.caon@epfl.ch; david.atienza@epfl.ch?subject=Visual X-HEEP: Enhancing X-HEEP with Camera Interfaces and VGA monitors'>davide.schiavone@epfl.ch; michele.caon@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project755minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Michele Caon, Prof. David Atienza<br>\";<\/script>\n<span id=project755><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Michele Caon, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project755',project755); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor750><\/a><b><span style='font-size: 20px;'>Optimization of Transformer-based ensemble<\/b><br><script>var project750=\"<p>The goal of the project is to explore and optimize a novel hardware\/software co-optimization strategy for Transformer inference, based on model ensembles, shared-indexes codebook and parallel processing.<\/p><p>Transformer are state-of-the-art algorithms in AI, with applications ranging from object detection\/classification to machine translation. Nonetheless, their high storage and computational requirements poses a challenge in resource-constrained edge computing scenarios. In this context, codebook-based optimizations [1] greatly reduce workloads and memory footprints of Transformer models by imposing few admissible weight values, stored in a small table (codebook) for each layer, and representing large weight matrices with codebook indexes.<\/p><p>By introducing constraints on weight values, codebooking degrades, to a degree, the accuracy of models. To recover from accuracy losses, multiple model instances, with the same structure but different codebooks, can be employed (a strategy named ensembling in [2]). Crucially, instances can be optimized to employ different codebooks, but <em>with the same indexes<\/em>, hence minimally impacting memory requirements. Moreover, SIMD extensions such as the ARM Scalable Vector Extension (SVE) can be employed to leverage the data parallelism deriving from ensembling to greatly reduce run time. While this approach has been proven efficient for Transformer benchmarks, it also introduces new challenges. In particular, Transformer models rely on multiplications between very large weight matrices, causing inefficient memory interaction and poor data locality.<\/p><p>In this project the student will be asked to develop strategies to minimize these downsides. In more detail, he\/she will develop an efficient tiling strategy to improve inference-time the data locality for ensembled and codebooked transformers. Furthermore, multiple tiles will be deployed on multiple computing units, exploring the benefits of a multicore approach. The student will be asked to perform critical analysis on the obtained results, evaluating the achieved speedup and identifying new emerging bottlenecks.<\/p><p>The project will be carried out at the <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded Systems Laboratory (ESL)<\/a> of EPFL. ESL comprises more than 40 researchers, active in many research in the hardware\/software co-design spectrum. The student will be under the supervision of and Mr. Stefano Albini, Dr. Giovanni Ansaloni, and Prof. David Atienza.<\/p><p>In order to achieve a passing grade of 4.0, the student is expected to:<\/p><ul><li>Become familiar with the existing codebase (including multiple codebooked Transformer models), learning the processing flow and the available SVE-accelerated transformer implementations.<\/li><li>Become familiar with the gem5 full system simulator, that will be used throughout the project to explore architectural parameters.<\/li><li>Implement an efficient tiling strategy for the matrix-multiplication at the core of the algorithm.<\/li><li>Implementation of a multi-core version of the tiling, distributing tiles to different computational units.<\/li><\/ul><p>Additionally, the tasks each add 1 additional point to the project grade:<\/p><ul><li>Extend the algorithm to work with different data representations (float16, float32)<\/li><li>Explore the run-time implications of alternative ways of sharing weights (between multiple heads of the same learner, groups of heads)<\/li><\/ul><h3><strong>Requirements:<\/strong><\/h3><ul><li>Proficiency in C language and modular programming.<\/li><li>Understanding of computer architecture, including cache and main memory interaction.<\/li><li>Basic command of Python language.<\/li><li>Familiarity with the git version control system.<\/li><li>Interest in AI optimization and acceleration.<\/li><\/ul><p>[1] Flavio Ponzina et al. &ldquo;Using ensemble learning to improve radiation tolerance of CNNs in space applications&rdquo;. In: Proceedings of SPAICE2024: The First Joint European Space Agency\/IAA Conference on AI in and for Space. 2024, pp. 16&ndash;20.<\/p><p>[2] Flavio Ponzina et al. &ldquo;An Accuracy-Driven Compression Methodology to Derive Efficient Codebook-Based CNNs&rdquo;. In: 2022 IEEE International Conference on Omni-layer Intelligent Systems (COINS). 2022, pp. 1&ndash;6. DOI: 10.1109\/COINS54846.2022.9854986.<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Stefano Albini, Dr. Giovanni Ansaloni, Prof. David Atienza<br> Contact email: <a href='mailto:stefano.albini@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch?subject=Optimization of Transformer-based ensemble'>stefano.albini@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project750minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Stefano Albini, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project750><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Stefano Albini, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project750',project750); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor749><\/a><b><span style='font-size: 20px;'>ML-based Design Space Exploration Framework for LLM Inference in Mixed-Precision Systolic Arrays<\/b><br><script>var project749=\" \t\t <p>This project focuses on the analysis, parameterization, validation, and optimization of the TiC-SAT systolic array architecture. The student will first study the TiC-SAT architecture, with particular attention to the processing element design, dataflow, and interconnect structure. Based on this understanding, the student will develop an evaluation framework to validate the hardware design and assess different architectural configurations.<\/p> <p>A systolic array consists of a network of tightly coupled data-processing units, referred to as Processing Elements (PEs), through which data flows in a structured manner. Each PE performs a small portion of the overall computation&mdash;typically multiply-accumulate (MAC) operations&mdash;and forwards intermediate results to neighboring elements. This architecture enables massive parallelism and high data reuse. By significantly reducing accesses to off-chip memory, systolic arrays achieve exceptional energy efficiency for matrix-intensive workloads, which form the computational backbone of deep learning models and large language models (LLMs).<\/p> <p>The project will extend the existing TiC-SAT design to support multiple systolic array sizes and numerical precision configurations. These parameterized designs will then be evaluated in terms of hardware and application-level metrics, including latency, area, and perplexity. Finally, the student will investigate machine-learning-based design space exploration methods to efficiently identify Pareto-optimal configurations.<\/p> <p>Projects tasks will hence be at the crossroad of hardware design and machine learning. In particular, machine learning will be involved in two aspects of the project. On the benchmarking side, the systolic array will be evaluated using Llama-1B as a representative machine learning workload to assess its performance. On the design exploration side, machine learning will be used to guide the search through the design space. Possible approaches include reinforcement learning, Bayesian optimization, active learning, or regression-based performance prediction.<\/p> <p><strong>Tasks description<\/strong><\/p> <ol> <li>Understand the overall TiC-SAT systolic array architecture, including the processing element design, dataflow, interconnect structure, and supported precision formats.<\/li> <li>Extend or adapt the scalable PE template to generate systolic array designs with different array sizes and several bitwidth configurations in the same instance.<\/li> <li>Prepare a testbench to verify the correctness of the TiC-SAT processing element and systolic array behavior. The framework should support functional validation and performance evaluation across different architectural configurations.<\/li> <li>Assess the generated designs using relevant metrics such as latency, area, and model perplexity.<\/li> <li>Implement at least one machine-learning-based design space exploration method to guide the search for efficient architecture configurations.<\/li> <li>Identify and report a set of Pareto-optimal solutions that trade off latency, area, and perplexity.<\/li> <\/ol> <p><strong>Project objectives<br \/><\/strong>The fulfillment of the following mandatory tasks will result in a passing grade of 4.0:<\/p> <ul> <li>[Mandatory] A testbench for testing and validating the TiC-SAT processing element design. (1 pt.)<\/li> <li>[Mandatory] A parameterized systolic array architecture supporting different array sizes and different numerical precision configurations inside the same instance. (0.5 pt.)<\/li> <li>[Mandatory] An evaluation framework for measuring latency, area, and perplexity across different configurations. (0.5 pt.)<\/li> <li>[Mandatory] A basic machine-learning-based design space exploration method for navigating the architecture design space, exploring at least 5 design solutions in terms of latency and perplexity. (1 pt.)<\/li> <\/ul> <p>Furthermore, the fulfillment of the following two optional tasks will result in 2 additional points up to the 6.0 maximum grade:<\/p> <ul> <li>[Optional] An multi-objective machine-learning-based design space exploration method for navigating the architecture design space, navigating more than 20 design solutions in terms of latency, area, and perplexity. (1 pt.)<\/li> <li>[Optional] Implement and compare multiple ML-based design space exploration methods, studying the trade-off between design space exploration time and the optimality of the selected solutions. (1 pt.)<\/li> <\/ul> <p><strong>Required knowledge and skills<\/strong><\/p> <ul> <li>Proficiency in RTL design and programming (e.g., VHDL or Verilog).<\/li> <li>Basic understanding of computer architecture.<\/li> <li>Strong analytical thinking and scientific curiosity.<\/li> <\/ul> <p><strong>References<\/strong><\/p> <p>[1] A. Amirshahi, J. Klein, G. Ansaloni, D. Atienza, &ldquo;TiC-SAT: Tightly-coupled Systolic Accelerator for Transformers&rdquo;,  28th Asia and South Pacific Design Automation Conference (ASP-DAC &rsquo;23), Tokyo, Japan, doi: 10.1145\/3566097.3567867<\/p> <p>[2] S. Machetti, P. D. Schiavone, G. Ansaloni, M. Pe&oacute;n-Quir&oacute;s and D. Atienza, &ldquo;X-HEEP: An Open-Source, Configurable and Extendible RISC-V Platform for TinyAI Applications,&rdquo; <em>2025 IEEE Computer Society Annual Symposium on VLSI (ISVLSI)<\/em>, Kalamata, Greece, 2025, pp. 1-6, doi: 10.1109\/ISVLSI65124.2025.11130281.<\/p> <p><strong>Type of work<\/strong><\/p> <ul> <li><strong>40%<\/strong> <strong>HW design<\/strong>: Development and implementation of the parameterized architecture, including mixed-bitwidth systolic array hardware.<\/li> <li><strong>40% Design Space Exploration<\/strong>: Evaluation of ML-based design space exploration methods for selecting hardware design parameters.<\/li> <li><strong>20% Performance Evaluation<\/strong>: Benchmark and analyze representative machine learning workloads on the proposed hardware design.<\/li> <\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> \";<\/script>\n<script>var project749minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project749><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Ms. Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project749',project749); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td>\n    <div style='position:relative;'>\n     <div style='position: absolute;top:-80px;left:-300px;'>\n       <img border=0 src=https:\/\/eslweb.epfl.ch\/projects\/images\/notavailable.gif alt='project no longer available'>\n     <\/div>\n    <\/div><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor748><\/a><b><span style='font-size: 20px;'>Comparative Evaluation of Time-Series Embedding Methods for Seizure Classification<\/b><br><script>var project748=\"<p>Wearable sensors and mobile devices continuously generate large volumes of time-series data. These signals enable applications such as Human Activity Recognition (HAR) and accelerometer-based seizure detection. A central challenge in both domains is learning representations that capture meaningful temporal patterns while remaining robust to noise, inter-subject variability, and domain shifts.<\/p> <p>Traditional approaches like convolutional neural networks learn task-specific features, while newer self-supervised models like <strong>TS2Vec<\/strong> learn universal representations that can generalize to multiple downstream tasks. Additionally, handcrafted features based on domain knowledge remain competitive in many settings, especially when data is limited.<\/p> <p>This project will explore and compare three distinct embedding strategies (<strong>CNN-based end-to-end models<\/strong>, <strong>TS2Vec self-supervised embeddings<\/strong>, and <strong>manual feature extraction<\/strong>) focusing on their classification accuracy and representational quality across multiple HAR datasets and a seizure dataset.<\/p> <h2><strong>Objectives<\/strong><\/h2> <ol> <li>Implement and compare three embedding approaches:<\/li> <\/ol> <p>\u25cb  A baseline CNN trained end-to-end.<\/p> <p>\u25cb  Self-supervised TS2Vec embeddings with downstream classifiers.<\/p> <p>\u25cb  Manual feature extraction with classical machine learning models.<\/p> <ol> <li>Evaluate classification performance across multiple HAR datasets (e.g., Capture-24, WISDM, Opportunity, recgym) and a private seizure dataset.<\/li> <li>Analyse the quality of learned representations using clustering and visualization methods.<\/li> <li>Assess cross-domain generalization, particularly transfer from HAR data to seizure classification.<\/li> <li>(Optional) Investigate feasibility for deployment on resource-constrained edge devices.<\/li> <\/ol> <h2><strong>Evaluation Metrics<\/strong><\/h2> <ul> <li><strong>Classification:<\/strong> Accuracy, precision\/recall, F1-score, ROC-AUC, false positive rate<\/li> <li><strong>Representation quality:<\/strong><\/li> <\/ul> <p>\u25cb  Dimensionality reduction visualisations (e.g., t-SNE, UMAP).<\/p> <p>\u25cb  Clustering metrics (e.g., silhouette score).<\/p> <h3><strong>Grading Criteria and Task Distribution<\/strong><\/h3> <h4><u>Minimum Requirements (Grade: 4.0)<\/u><\/h4> <p>To obtain a passing grade, the student must successfully complete the following core tasks:<\/p> <ul> <li><strong>Implementation<\/strong>: Develop and validate three baseline models: Supervised CNN (end-to-end), TS2Vec (self-supervised), and manual feature extraction.<\/li> <li><strong>Data Processing<\/strong>: Standardize and evaluate the models on at least two HAR datasets and the provided seizure dataset.<\/li> <li><strong>Evaluation<\/strong>: Report standard classification metrics, including Accuracy, F1-score, and ROC-AUC.<\/li> <li><strong>Reporting<\/strong>: Submit a report\/presentation explaining the methodology, experimental setup, and a basic comparative analysis of results.<\/li> <\/ul> <h4><u>Optional Tasks for Higher Grades (Up to 6.0)<\/u><\/h4> <p>The final grade increases by approximately 0.5 points for each additional task completed, provided the execution meets professional standards:<\/p> <ul> <li><strong>Representational Analysis: <\/strong>Quantify embedding quality using dimensionality reduction (t-SNE\/UMAP) and clustering metrics (e.g., silhouette scores).<\/li> <li><strong>Cross-Domain Generalization: <\/strong>Evaluate transfer learning performance by applying representations learned from HAR datasets to seizure detection tasks.<\/li> <li><strong>Extended Benchmarking: <\/strong>Scale the evaluation to include the full suite of six HAR datasets (Capture-24, WEAR, WISDM, Opportunity, and Recgym).<\/li> <li><strong>Edge Deployment Study: <\/strong>Profile computational latency and memory usage, and implement a C-based inference prototype for resource-constrained hardware.<\/li> <\/ul> <p>&nbsp;<\/p> <h2><strong>Requirements<\/strong><\/h2> <ul> <li>Basic understanding of signal processing and time-series analysis.<\/li> <li>Good programming skills in Python.<\/li> <li>Experience with machine learning tools (PyTorch, scikit-learn).<\/li> <li>(Optional) C programming and firmware development<\/li> <\/ul> <p><strong>Type of Work<\/strong><\/p> <ul> <li><strong>33%<\/strong>: Literature review, methodological design, representation analysis, and interpretation of results.<\/li> <li><strong>67%<\/strong>: Implementation of models and experimental evaluation in Python<\/li> <\/ul> <p><strong>Sources<\/strong><\/p> <ul> <li><a href='https:\/\/github.com\/zhihanyue\/ts2vec'>Ts2vec repo<\/a><\/li> <li><a href='https:\/\/github.com\/ElsevierSoftwareX\/SOFTX%5F2020%5F1?tab=readme-ov-file#time-series-feature-extraction-library'>Time Series Feature Extraction Library<\/a><\/li> <li>Human activity recognition datasets:<br \/>&ndash; capture24: <a href='https:\/\/ora.ox.ac.uk\/objects\/uuid:99d7c092-d865-4a19-b096-cc16440cd001'>https:\/\/ora.ox.ac.uk\/objects\/uuid:99d7c092-d865-4a19-b096-cc16440cd001<br \/><\/a> &ndash; WEAR: <a href='https:\/\/github.com\/drhashimali\/wear'>https:\/\/github.com\/drhashimali\/wear<br \/><\/a> &ndash; HARTH: <a href='https:\/\/archive.ics.uci.edu\/dataset\/779\/harth'>https:\/\/archive.ics.uci.edu\/dataset\/779\/harth<br \/><\/a> &ndash; WISDM: <a href='https:\/\/www.cis.fordham.edu\/wisdm\/dataset.php'>https:\/\/www.cis.fordham.edu\/wisdm\/dataset.php<br \/><\/a> &ndash; Opportunity: <a href='https:\/\/archive.ics.uci.edu\/dataset\/226\/opportunity+activity+recognition'>https:\/\/archive.ics.uci.edu\/dataset\/226\/opportunity+activity+recognition<br \/><\/a> &ndash; Recgym: <a href='https:\/\/zhaxidele.github.io\/RecGym\/'>https:\/\/zhaxidele.github.io\/RecGym\/<\/a><\/li> <\/ul> \t<br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dimitra Tatlin, Dr. Lara Orlandic, Dr. Jonathan Dan, Prof. David Atienza<br> Contact email: <a href='mailto:lara.orlandic@epfl.ch;jonathan.dan@epfl.ch;david.atienza@epfl.ch;dimitra.tatli@epfl.ch?subject=Comparative Evaluation of Time-Series Embedding Methods for Seizure Classification'>lara.orlandic@epfl.ch;jonathan.dan@epfl.ch;david.atienza@epfl.ch;dimitra.tatli@epfl.ch<\/a><br>\";<\/script>\n<script>var project748minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dimitra Tatlin, Dr. Lara Orlandic, Dr. Jonathan Dan, Prof. David Atienza<br>\";<\/script>\n<span id=project748><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dimitra Tatlin, Dr. Lara Orlandic, Dr. Jonathan Dan, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project748',project748); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor746><\/a><b><span style='font-size: 20px;'>Precision-Scalable Systolic Arrays for High-Efficiency LLM Inference Acceleration<\/b><br><script>var project746=\"<p>This project focuses on the design and implementation of a scalable-bitwidth systolic array, a specialized hardware accelerator aimed at improving the computational efficiency of modern artificial intelligence workloads. Its primary objective is to develop an architecture that can dynamically support varying numerical precisions through a template-based design, ranging from low-bit quantization (e.g., INT2\/INT3\/INT4) to higher-precision formats such as FP16.<\/p>    <p>A systolic array consists of a network of tightly coupled Processing Elements (PEs), through which data flows in a structured manner. Each PE performs a small portion of the overall computation&mdash;typically multiply-accumulate (MAC) operations&mdash;and forwards intermediate results to neighboring elements. Systolic arrays enables massive parallelism and high data reuse. By significantly reducing accesses to off-chip memory, they achieve very high energy efficiency for matrix multiplication workloads, which form the computational backbone of deep learning and large language models.<\/p>    <p>Despite their efficiency, state-of-the-art systolic arrays are typically limited to a single, fixed bitwidth. Most existing designs target a specific numerical format, such as INT8 or FP16, which makes them less adaptable to the diverse precision requirements of modern workloads. As a result, these architectures struggle to efficiently support varying numerical characteristics, limiting their flexibility and broader applicability.<\/p>    <p>To address this limitation, this project proposes a reconfigurable, scalable-bitwidth systolic array. To this end, the student will be tasked with the development of a template-based PE architecture that enables the generation of practical systolic arrays supporting varying bitwidths. Multiple PEs can be dynamically grouped to form scalable matrix-multiplication accelerators with configurable precision. This approach aims to bridge the gap between flexibility and efficiency, enabling hardware accelerators that better match the diverse precision demands of modern AI models.<\/p>    <p><strong>Tasks description<\/strong><\/p>    <ol class='wp-block-list'> <li>Understand the architecture of the TiC-SAT systolic array under development at ESL_EPFL, including its PE design and interconnects.<\/li>    <li>Implement a PE design template that supports various bitwidth configuration.<\/li>    <li>Generate and test heterogeneous systolic array hardware designs using the scalable-bitwidth PEs to enable flexible precision support.<\/li> <\/ol>    <p><strong>Project objectives<br \/><\/strong>The fulfillment of the following objective is required for a passing grade (4.0)<\/p>    <ul class='wp-block-list'> <li>Extend the TiC-SAT PE design to support multiple bitwidth, including a testbench and testsuite to validate the design.<\/li>    <li>Create a template-based generator of the systolic array. The template must enable the generation of instances of the systolic array from configuration parameters, specifying the array size, supported &nbsp;bitwidths and the arrangement of PEs.<\/li>    <li>Characterize the runtime latency and energy efficiency of the modified systolic array for at least four different bitwidth configurations.<\/li> <\/ul>    <p>The completion of each of the following tasks will add 1.0 extra points to the project grade<\/p>    <ul class='wp-block-list'> <li>Build a functional simulator for the SA in python.<\/li>    <li>Emulate the system in an FPGA.<\/li> <\/ul>    <p><strong>Required knowledge and skills<\/strong><\/p>    <ul class='wp-block-list'> <li>Proficiency in RTL design and programming (VHDL or Verilog).<\/li>    <li>Basic understanding of computer architecture.<\/li>    <li>Strong analytical thinking and scientific curiosity.<\/li> <\/ul>    <p><strong>References<\/strong><\/p>    <p>[1] A. Amirshahi, J. Klein, G. Ansaloni, D. Atienza, &ldquo;TiC-SAT: Tightly-coupled Systolic Accelerator for Transformers&rdquo;,&nbsp; 28th Asia and South Pacific Design Automation Conference (ASP-DAC &rsquo;23), Tokyo, Japan, doi: 10.1145\/3566097.3567867<\/p>    <p>[2] S. Machetti, P. D. Schiavone, G. Ansaloni, M. Pe&oacute;n-Quir&oacute;s and D. Atienza, &ldquo;X-HEEP: An Open-Source, Configurable and Extendible RISC-V Platform for TinyAI Applications,&rdquo;&nbsp;<em>2025 IEEE Computer Society Annual Symposium on VLSI (ISVLSI)<\/em>, Kalamata, Greece, 2025, pp. 1-6, doi: 10.1109\/ISVLSI65124.2025.11130281.<\/p>    <p><strong>Type of work<\/strong><\/p>    <ul class='wp-block-list'> <li>80% <strong>HW design<\/strong>: development and implementation of the scalable-bitwidth systolic array hardware.<\/li>    <li>20% <strong>Performance Evaluation<\/strong>: Benchmarking and analysis of machine learning workloads on the designed hardware.<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> Contact email: <a href='mailto:ruben.rodriguezalvarez@epfl.ch; yuxuan.wang@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch?subject=Precision-Scalable Systolic Arrays for High-Efficiency LLM Inference Acceleration'>ruben.rodriguezalvarez@epfl.ch; yuxuan.wang@epfl.ch; giovanni.ansaloni@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project746minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br>\";<\/script>\n<span id=project746><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Rub\u00e9n Rodr\u00edguez \u00c1lvarez, Yuxuan Wang, Dr. Giovanni Ansaloni, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project746',project746); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor741><\/a><b><span style='font-size: 20px;'>Development of a TFT display interface for VersaSens<\/b><br><script>var project741=\"<p>The project aims to develop a user interface to be implemented within the VersaSens sensing platform. VersaSens, developed at the Embedded Systems Laboratory (ESL), is a modular, multimodal, extendable, and reconfigurable Edge AI platform. It enables the integration of add-on modules alongside its array of sensing and processing modules, providing a flexible foundation for diverse applications.<\/p> <p>The main aim of this project is to develop a PCB module with an embedded touchscreen TFT display compatible with the VersaSens main board. A secondary aim is to develop firmware to visualize VersaSens-collected signals in real time. The touchscreen capabilities of the interface should be used to enhance the platform's user-friendliness.<\/p> <p>This project is linked to an ongoing research focus in ESL, targeting the development of a multimodal sensing platform with applications in the medical field. Thus, some aspects of this project will benefit from previous developments within the laboratory (Firmware, algorithms for physiological data processing with embedded deployment), and it is expected that this semester's project will further improve and adapt the current solution.<\/p> <p><strong>Mandatory tasks<\/strong><\/p> <p>Completion of <strong>all<\/strong> these tasks is required to pass the exam and obtain a grade of <strong>4<\/strong>. Failure to complete any of these tasks will result in <strong>no pass<\/strong>:<\/p> <ol><li>Get familiar with VersaSens sensing platform. Conduct a comprehensive review of commercial TFT displays, including their specifications, and assess their compatibility with the VersaSens platform. Select a display according to VersaSens requirements and provide a justification.<\/li><li>Make the PCB design for the TFT display to be compatible with the VersaSens platform, and to fulfill the mechanical requirements of our casings. Prepare appropriately the manufacturing files and send them for production.<\/li><li>Develop the firmware for the communication between the microcontroller and the TFT display and integrate it to the VersaSens firmware. Validate the firmware by plotting real-time signals acquired from one of the sensors on screen (ECG, EEG, PPG, EMG, SKT, Sound, Bio-Z).<\/li><li>Further develop the firmware to activate the touchscreen capabilities of the TFT display.<\/li><li>Develop a user interface with a menu to access the plotting of all different sensors in real time (ECG, EEG, PPG, EMG, SKT, Sound, Bio-Z). The battery level should also be shown.<\/li><li>Measure the power consumption of the touchscreen and adapt the visualization for optimized consumption while keeping a user-friendly interface. Here are some potential parameters to assess: backlight brightness, refresh rate, and data processing for plotting...<\/li><li>Deliver a&nbsp;complete documentation package to be uploaded to the VersaSens GitHub repository, including all datasets, software, and firmware.<\/li><\/ol> <p><strong>Optional tasks<\/strong><\/p> <p>Once all mandatory tasks have been completed and a grade of <strong>4<\/strong> has been obtained, each optional task completed will contribute an additional 1 point to the final grade, up to a maximum grade of <strong>6<\/strong>:<\/p> <ol><li>Create a touchscreen keyboard to write a text to be saved on the SD card.<\/li><li>Create a menu to access and modify settings and parameters of the sensors in the firmware.<\/li><\/ol> <p><strong>Type of work<\/strong><\/p> <ul><li>40% firmware development<\/li><li>30% PCB development<\/li><li>25% Performance evaluation through systematic testing of the platform.<\/li><li>5% Preparation and delivery of the complete documentation package.<\/li><\/ul> <p><strong>Desired skills:<\/strong><\/p> <ul><li>Strong analytical skills<\/li><li>Experience in embedded systems development<\/li><li>Teamwork and git<\/li><\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. J\u00e9r\u00f4me Thevenot, Dr. Taraneh Aminosharieh Najafi, Doctoral candidate Wensi Zhang, Prof. David Atienza<br> Contact email: <a href='mailto:jerome.thevenot@epfl.ch;taraneh.aminoshariehnajafi@epfl.ch;wensi.zhang@epfl.ch;david.atienza@epfl.ch?subject=Development of a TFT display interface for VersaSens'>jerome.thevenot@epfl.ch;taraneh.aminoshariehnajafi@epfl.ch;wensi.zhang@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project741minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. J\u00e9r\u00f4me Thevenot, Dr. Taraneh Aminosharieh Najafi, Doctoral candidate Wensi Zhang, Prof. David Atienza<br>\";<\/script>\n<span id=project741><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. J\u00e9r\u00f4me Thevenot, Dr. Taraneh Aminosharieh Najafi, Doctoral candidate Wensi Zhang, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project741',project741); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor733><\/a><b><span style='font-size: 20px;'>Architecture Definition and FPGA Deployment of a Near-Memory Computing Platform for Edge AI and Multimedia<\/b><br><script>var project733=\"<p>The shift toward data-centric algorithms, particularly in Artificial Intelligence (AI), Machine Learning (ML), and multimedia processing, has exposed the &quot;memory wall&quot; bottleneck in traditional von Neumann architectures. In these systems, the energy cost of moving data between memory and the CPU can be up to 100x higher than the arithmetic operations themselves. Near-Memory Computing (NMC) addresses this by performing computation directly where data is stored, significantly reducing data movement and exploiting internal memory bandwidth.<\/p> <p>However, existing CIM solutions often require high implementation effort and lack the flexibility required for general-purpose software integration. To address this, the Embedded Systems Laboratory (ESL) has developed <a href='https:\/\/ieeexplore.ieee.org\/document\/10964076'><strong>NM-Carus<\/strong><\/a>, a fully-autonomous, vector-capable RISC-V programmable NMC unit. NM-Carus is designed as a drop-in replacement for traditional SRAM banks, offering a functionally transparent memory mode alongside a high-performance computing mode. By integrating multiple NM-Carus instances into the <a href='https:\/\/github.com\/x-heep\/x-heep'><strong>X-HEEP<\/strong><\/a> microcontroller platform, this project aims to create a scalable SoC capable of accelerating complex workloads like Visual Transformers and multimedia codecs with peak energy efficiencies.<\/p> <h3><strong>Project Description and Main Goal<\/strong><\/h3> <p>The primary objective of this project is to define, implement, and deploy an NMC-enhanced System-on-Chip (SoC) based on X-HEEP, a configurable, ultra-low-power RISC-V microcontroller developed at EPFL. The student will re-architect the X-HEEP memory subsystem by replacing traditional SRAM banks with instances of the NM-Carus IP. The resulting platform must be tailored to accelerate specific target applications, including multimedia libraries (e.g., image and video codecs) and AI models (e.g., visual transformer), requiring a deep exploration of hardware-software trade-offs.<\/p> <h3><strong>Project Objectives<\/strong><\/h3> <p>The project will require a systematic approach involving RTL design, verification, and FPGA emulation. The student will work in close collaboration with another candidate taking care of deploying a set of relevant applications on the developed SoC. The specific tasks are defined in the following sections.<\/p> <ul><li><strong>X-HEEP Platform Familiarization<\/strong><\/li><\/ul> <p>The student must first gain a comprehensive understanding of the X-HEEP architecture. This involves setting up the development environment, installing required EDA tools (e.g., Verilator) and software toolchains. The student is expected to follow X-HEEP documentation to learn how to build the simulation model, compile example software applications, and run RTL simulations. It is crucial for the student to understand the SoC's internal mechanisms, specifically how parameters modify the hardware generation, how peripherals are exposed to the host CPU and what features they offer, how X-HEEP can be extended with new accelerators or coprocessors, and how interrupts are handled. Following simulation, the student will synthesize the design for FPGA and cooperate with the software team to deploy baseline versions of the benchmark applications, establishing a performance reference.<\/p> <ul><li><strong>NM-Carus Architecture Familiarization<\/strong><\/li><\/ul> <p>The student needs to acquire in-depth knowledge of the NM-Carus device, including its hardware microarchitecture and its custom RISC-V ISA extension. Understanding the vector processing capabilities, memory management paradigm, and the interface mechanism with the host system is essential for later optimization of processing kernels.<\/p> <ul><li><strong>Architecture Definition and Integration<\/strong><\/li><\/ul> <p>Starting from a provided template containing initial NM-Carus instances, the student will define the NMC-enhanced architecture. This is a critical design phase where the student must investigate the optimal configuration, bus architecture, and the number and size of the NMC devices to best suit the expected workloads. Key design decisions will include the bus architecture (e.g., OBI\/AHB crossbars), memory layout (interleaving schemes, banking factors), and the mapping of system addresses to the NM-Carus Vector Register Files (VRF). The student must also design efficient communication mechanisms between the host CPU and the NMC instances, as well as inter-NMC communication if deemed necessary.<\/p> <ul><li><strong>Functional Verification<\/strong><\/li><\/ul> <p>Once the base platform is defined, the student will verify its functionality through RTL simulation and FPGA emulation. This involves deploying simple &quot;sanity check&quot; applications (e.g. matrix multiplication, data movement tests), developed in collaboration with the software team, to ensure the modified memory subsystem functions correctly as both standard memory and a computing unit. The student will modify these initial tests to exploit the parallelism of multiple NMC instances.<\/p> <ul><li><strong>Benchmark Deployment and Profiling<\/strong><\/li><\/ul> <p>The student will assist the software team in deploying the full target benchmarks. This task involves hardware-level profiling to identify bottlenecks in the computing kernels. The student will explore opportunities to extend the hardware, such as adding custom instructions to the NM-Carus ISA, modifying the banking structure to increase bandwidth, or adding system-level modules to minimize data transfer latency and streamline synchronization with the host core and between NM-Carus instances.<\/p> <ul><li><strong>Architectural Design Space Exploration<\/strong><\/li><\/ul> <p>Based on profiling results, the student will iterate on the architecture. This may involve changing the number and size of NM-Carus instances, adjusting internal banking parallelism, or refining the interconnect policy to maximize system throughput and energy efficiency for the selected workloads.<\/p> <ul><li><strong><em>[Optional] <\/em><\/strong><strong>Physical Implementation<\/strong><\/li><\/ul> <p>Implement the SoC on a 65nm technology node (logic synthesis) to extract timing and energy consumption figures via post-synthesis simulations.<\/p> <h3><strong>Working Environment<\/strong><\/h3> <p>The research will take place at the Embedded Systems Laboratory (ESL) at EPFL, a globally recognized institution for research in embedded systems and computer architecture. ESL offers a stimulating and collaborative research environment, complete with access to cutting-edge tools and resources. The candidate will have the chance to work closely with Prof. David Atienza and other members of the ESL team. The candidate is expected to work in close collaboration with another student taking care of the software aspects of the project.<\/p> <h3><strong>Expected Outcomes and Impact<\/strong><\/h3> <p>The successful completion of this project will result in:<\/p> <ul><li>A novel, scalable SoC architecture exploiting programmable near-memory computing IPs to optimize the execution of multimedia and AI tasks on edge devices.<\/li><li>A functional FPGA prototype of such architecture.<\/li><li>A detailed performance assessment of the proposed architecture while running a predefined set of benchmarks applications.<\/li><\/ul> <h3><strong>Prerequisites<\/strong><\/h3> <ul><li>Strong background in computer architecture and digital systems design.<\/li><li>In-depth knowledge of the RISC-V architecture (ISA and microarchitecture).<\/li><li>Proficiency in SystemVerilog\/Verilog for RTL design and verification.<\/li><li>Experience with FPGA synthesis and emulation flows (e.g., Xilinx Vivado).<\/li><li>Familiarity with low-level programming (C and RISC-V assembly) to understand hardware-software interactions.<\/li><li>Experience with version control systems (Git) and Linux-based development environments.<\/li><li>Advanced experience with collaborative software and hardware development using Git.<\/li><\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul><li>Knowledge of bus protocols (AHB, AXI, OBI).&nbsp;<\/li><li>Familiarity with logic synthesis tools (Synopsys Design Compiler or similar).<\/li><li>Advanced proficiency in English.&nbsp;&nbsp;<\/li><li>Effective communication skills.<\/li><\/ul> <p><strong>Type of work<\/strong><\/p> <ul><li>70% hardware design, RTL implementation, and FPGA emulation.<\/li><li>15% collaboration on software deployment and hardware-software co-design.<\/li><li>15% architecture analysis, documentation, and reporting.<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Michele Caon, Dr. Davide Schiavone, Prof. David Atienza<br> Contact email: <a href='mailto:michele.caon@epfl.ch;davide.schiavone@epfl.ch;david.atienza@epfl.ch?subject=Architecture Definition and FPGA Deployment of a Near-Memory Computing Platform for Edge AI and Multimedia'>michele.caon@epfl.ch;davide.schiavone@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project733minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Michele Caon, Dr. Davide Schiavone, Prof. David Atienza<br>\";<\/script>\n<span id=project733><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Michele Caon, Dr. Davide Schiavone, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project733',project733); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor732><\/a><b><span style='font-size: 20px;'>Application Deployment and Software Optimization on a Near-Memory Computing Platform for Edge AI and Multimedia<\/b><br><script>var project732=\"<p>While Near-Memory Computing (NMC) hardware offers significant theoretical gains in energy efficiency and throughput, unlocking this potential requires a specialized software stack and tailored application mapping. Traditional embedded software often struggles with the &quot;memory wall,&quot; where data movement dominates the power budget. This project focuses on porting and optimizing real-world, data-intensive applications (specifically transformer neural networks and multimedia codecs) onto a novel NMC-enhanced SoC.<\/p> <p>The target platform integrates the <a href='https:\/\/github.com\/x-heep\/x-heep'><strong>X-HEEP<\/strong><\/a> microcontroller with <a href='https:\/\/ieeexplore.ieee.org\/document\/10964076'><strong>NM-Carus<\/strong><\/a> units. NM-Carus provides a software-programmable solution based on the RISC-V ISA, allowing it to function as either a standard SRAM bank or a vector processing unit. The student will develop the firmware, drivers, and optimized kernels necessary to transition from standard CPU execution to accelerated NMC-based execution on a multi-instance, multiple-instruction, multiple-data architecture, aiming to significantly reduce the system-level energy footprint.<\/p> <h3><strong>Project description and main goal<\/strong><\/h3> <p>The primary objective of this project is to deploy, profile, and optimize a set of industry-relevant benchmark applications&mdash;multimedia libraries (libpng, ffmpeg) and AI models (dinov2)&mdash;on an X-HEEP microcontroller enhanced with NM-Carus devices in place of conventional SRAM banks in its memory subsystem. The student will be responsible for creating a performance baseline on a standard RISC-V microcontroller, porting the applications to the NMC-enhanced variant, and rewriting critical computing kernels to exploit the near-memory vector capabilities.<\/p> <h3><strong>Project objectives<\/strong><\/h3> <p>The project will involve a mix of embedded software development, algorithm analysis, and hardware-software co-design. The student will work in close collaboration with another candidate taking care of developing the SoC and deploying it on FPGA for faster prototyping. The specific tasks are defined in the following sections.<\/p> <ul><li><strong>X-HEEP Platform Familiarization<\/strong><\/li><\/ul> <p>The student will start by mastering the X-HEEP software development environment. This involves understanding the CMake-based build system, the linker scripts, and the Hardware Abstraction Layer (HAL). Special focus must be placed on how peripherals (e.g., UART, DMA, Timers) are mapped in memory and exposed to the software, and how to compile, flash, and debug bare-metal RISC-V applications.<\/p> <ul><li><strong>NM-Carus ISA Familiarization<\/strong><\/li><\/ul> <p>The student must gain a deep understanding of the NM-Carus programming model. This includes studying its custom RISC-V ISA extensions, the distinction between its &quot;standard memory&quot; and &quot;computing&quot; modes, and the cycle-accurate performance of its vector instructions. Understanding the theoretical peak throughput is crucial for setting optimization targets and optimizing the processing kernels in later stages of the project.<\/p> <ul><li><strong>Benchmark Analysis<\/strong><\/li><\/ul> <p>Before moving to the novel SoC, the student will analyze the selected benchmarks (libpng, ffmpeg, dinov2) on a standard host PC. The goal is to understand the algorithmic flow, data structures, and dependencies of these complex libraries. The student will identify the &quot;hotspots&quot; that are prime candidates for acceleration (e.g., vectorizable functions like convolution or matrix multiplication).<\/p> <ul><li><strong>Baseline Porting and Profiling<\/strong><\/li><\/ul> <p>The student will port the selected applications to the standard X-HEEP platform. This is a non-trivial task that involves removing OS dependencies (e.g., filesystem calls), managing limited memory resources, and compiling for a bare-metal RISC-V target. Once ported, the student will profile the applications using hardware performance counters to establish a firm baseline for execution time, identifying the specific bottlenecks to be addressed.<\/p> <ul><li><strong>NMC Acceleration and Optimization<\/strong><\/li><\/ul> <p>Leveraging the analysis from the previous steps, the student will offload the identified critical kernels to the NM-Carus instances. This involves rewriting specific functions and offloading custom assembly kernels to the available NMC vector units in SIMD or MIMD fashion. The student will work in strict cooperation with the hardware team to optimize data layout (e.g., tiling data to fit into NM-Carus banks) and minimize the overhead of data transfers between the host CPU and the NMC units.<\/p> <ul><li><strong>Performance Evaluation and Design Iteration<\/strong><\/li><\/ul> <p>The NMC-accelerated applications will be profiled and compared against the baseline in terms of execution time. Based on the optimization difficulties or bottlenecks encountered during software development (e.g., need for a specific shuffle instruction, or better DMA synchronization), the student will provide feedback to the hardware team to trigger architectural improvements.<\/p> <h3><strong>Working environment<\/strong><\/h3> <p>The research will take place at the Embedded Systems Laboratory (ESL) at EPFL, a globally recognized institution for research in embedded systems and computer architecture. ESL offers a stimulating and collaborative research environment, complete with access to cutting-edge tools and resources. The candidate is expected to work in close collaboration with another student taking care of the hardware definition and implementation.<\/p> <h3><strong>Expected Outcomes and Impact<\/strong><\/h3> <p>The successful completion of this project will result in:<\/p> <ul><li>A set of optimized, real-world applications running on a novel near-memory computing SoC.<\/li><li>An efficient and flexible device driver and SDK for the NM-Carus NMC device.<\/li><li>A comprehensive performance assessment of the proposed SoC when running the selected benchmark applications.<\/li><\/ul> <h3><strong>Prerequisites<\/strong><\/h3> <ul><li>Strong background in computer architecture and embedded software.<\/li><li>Advanced proficiency in C\/C++ programming.<\/li><li>Proficient in low-level programming, ideally with RISC-V assembly.<\/li><li>Experience with cross-compilation toolchains (GCC\/LLVM) and Make\/CMake build systems.<\/li><li>Experience with bare-metal programming (no OS) and debugging (GDB).<\/li><li>Advanced experience with collaborative software and hardware development using Git.<\/li><li>Strong scripting and data collections and analysis skills.<\/li><li>Good analytical skills.<\/li><\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul><li>Knowledge of image processing algorithms, video compression standards, or neural network internals (specifically Transformers).<\/li><li>Advanced proficiency in English.&nbsp;&nbsp;<\/li><li>Effective communication skills.<\/li><\/ul> <p><strong>Type of work<\/strong><\/p> <ul><li>50% application software development and deployment.<\/li><li>35% hardware\/software co-design, implementation, verification, and validation.<\/li><li>15% profiling, documentation, and reporting.<\/li><\/ul><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Michele Caon, Dr. Davide Schiavone, Prof. David Atienza<br> Contact email: <a href='mailto:michele.caon@epfl.ch; davide.schiavone@epfl.ch; david.atienza@epfl.ch?subject=Application Deployment and Software Optimization on a Near-Memory Computing Platform for Edge AI and Multimedia'>michele.caon@epfl.ch; davide.schiavone@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project732minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Michele Caon, Dr. Davide Schiavone, Prof. David Atienza<br>\";<\/script>\n<span id=project732><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Michele Caon, Dr. Davide Schiavone, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project732',project732); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor710><\/a><b><span style='font-size: 20px;'>Mind the Gaps: Missing-Aware Embedding Layers for Transformer Forecasting<\/b><br><script>var project710=\"<div class='entry-content mb-5'> \t\t <p>Urban wastewater time-series collected from underground sensing  infrastructure are inherently noisy, irregular, and incomplete.  Communication interruptions, fluctuating battery levels, harsh  environmental conditions, and irregular sampling frequently lead to <em>missing timestamps<\/em>, <em>bursty gaps<\/em>, or <em>temporarily corrupted sequences<\/em>. Forecasting methods deployed in digital twins [2] must therefore handle incomplete temporal signals gracefully.<\/p> <p>Transformer architectures rely heavily on embedding layers that  convert raw time-series into a sequence of tokens. However, most  existing embedding mechanisms&mdash;such as convolutional patch embeddings or  positional encodings&mdash;assume <strong>dense, regularly sampled<\/strong> time-series.  When timestamps are missing, the embedding becomes misaligned,  positional encodings become unreliable, and attention mechanisms degrade  due to incomplete temporal structure. This loss of representational  quality reduces forecasting robustness, particularly for long horizons  or rainfall-driven surges.<\/p> <p>This semester project focuses on designing <strong>missing-aware embedding layers<\/strong>  specifically for hydrological time-series forecasting with  Transformers. The goal is to ensure that the model can ingest irregular  or incomplete sequences without degrading performance, thereby improving  reliability for real-world deployments of AquaCast [1] or related  forecasting systems.<\/p> <h3><strong>TASKS&nbsp;<\/strong><\/h3> <ol><li><strong> Literature Review<\/strong><ul><li>Conduct a literature review on embedding layers of Transformers for standard time-series and time-series with missing values.<\/li><\/ul><\/li> <li><strong> Dataset Preparation<\/strong><ul><li>Preprocess and construct the dataset from raw wastewater and precipitation records for this specific task.<\/li><li>Introduce controlled missingness (random, burst, structural) to evaluate robustness.<\/li><\/ul><\/li>  <li><strong> Embedding Layer Design, Implementation &amp; Testing<\/strong><ul><li>Design, implement, and test <strong>different embedding layers<\/strong> for a vanilla Transformer to improve forecasting while enabling the network to handle missing time-steps.<\/li><li>Implement the developed embedding layer for AquaCast, while considering both exogenous and endogenous time-series.<\/li><li>Provide quantitative measurements and interpretability for current methods and your proposed designs.<\/li><\/ul><\/li> <\/ol> <p><strong>Optional:<\/strong><\/p> <ul><li><strong>Missing samples representation:<\/strong> Investigate self-supervised reconstruction (e.g., masked-signal modeling) and evaluate transfer to forecasting tasks.<\/li><\/ul> <p><strong>REQUIREMENTS<\/strong><\/p> <ul><li>Solid Python programming, basic understanding of deep learning.<\/li><li>Interest in robust AI and time-series forecasting.<\/li><li>Scientific curiosity<\/li><\/ul> <h3><strong>TYPE OF WORK<\/strong><\/h3> <ul><li><strong>30% theory<\/strong> (study of embedding mechanisms, missing-data theory).<br \/><strong>70% implementation<\/strong> (embedding design, integration, benchmarking, analysis).<\/li><\/ul> <p><strong>REFERENCES<\/strong><\/p> <p>[1] Abdollahinejad, Golnoosh, Saleh Baghersalimi, Denisa-Andreea Constantinescu, Sergey Shevchik, and David Atienza. &ldquo;<strong>AquaCast<\/strong>: Urban Water Dynamics Forecasting with Precipitation-Informed Multi-Input Transformer.&rdquo; <a href='https:\/\/arxiv.org\/abs\/2509.09458'><em>https:\/\/arxiv.org\/abs\/2509.09458<\/em><\/a><\/p> <p>[2] UrbanTwin project: <a href='https:\/\/urbantwin.ch\/'>https:\/\/urbantwin.ch\/<\/a><\/p> \t<\/div>                         <div class='post-nav py-md-1'>                                 <div class='nav-prev'>         <\/div><\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br> Contact email: <a href='mailto: golnoosh.abdollahinejad@epfl.ch; denisa.constantinescu@epfl.ch; david.atienza@epfl.ch?subject=Mind the Gaps: Missing-Aware Embedding Layers for Transformer Forecasting'> golnoosh.abdollahinejad@epfl.ch; denisa.constantinescu@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project710minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br>\";<\/script>\n<span id=project710><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Golnoosh Abdollahinejad, Dr. Denisa Constantinescu, Prof. Dr. David Atienza<br> <a href=#_ onclick=opendesc('project710',project710); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/urbantwin.ch target=_blank title='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/197.png width=70 alt='An urban digital twin for climate action: Assessing policies and solutions for energy, water and infrastructure'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor660><\/a><b><span style='font-size: 20px;'>Extraction of Heart-Related Information in Epileptic Patients<\/b><br><script>var project660=\"<div class='entry-content'> \t\t\t <p>Epilepsy is a neurological condition that can be accompanied by  changes in heart rate and heart rate variability around seizure events.  Portable ECG (electrocardiogram) devices and smartwatch based PPG  (photoplethysmogram) sensors provide a unique opportunity to extract  heart-related features in real-time, enabling continuous monitoring for  patients with epilepsy. This project aims to leverage ECG and PPG data  from wearable devices to assess the quality of heart-related signals  before, during, and after seizures. Furthermore, the project will  replicate existing research on seizure detection using heart-related  features to investigate the feasibility of using wearable devices for  seizure detection.<\/p>    <p>The technical challenges include handling noisy ECG and PPG data,  ensuring signal quality in diverse patient conditions, and designing  algorithms capable of detecting relevant heart-related features during  seizure events. The system will need to analyze signal quality and  extract meaningful heart-related features while maintaining real-time  performance on resource-constrained devices like smartwatches.<\/p>    <p>This project will contribute to the field of epilepsy monitoring by  providing a method for using heart-related data from wearable devices to  assess seizure states. It could serve as a valuable tool for  researchers and clinicians working on seizure prediction and monitoring.<\/p>    <h3 class='wp-block-heading'>Tasks:<\/h3>    <ul class='wp-block-list'><li>Preprocess ECG and PPG signals from smartwatch data to remove noise and artifacts, especially during seizure events.<\/li><li>Implement algorithms to extract heart-related features from ECG and  PPG signals (e.g., heart rate, heart rate variability, and respiratory  rate).<\/li><li>Assess the quality of heart-related signals before, during, and  after seizures, identifying any characteristic changes in the features.<\/li><li>Replicate a seizure detection study using extracted heart-related  features to demonstrate the potential of using these features for  seizure detection.<\/li><\/ul>    <h3 class='wp-block-heading'>Requirements:<\/h3>    <ul class='wp-block-list'><li>Strong Python programming skills<\/li><li>Basic signal processing knowledge<\/li><li>Machine learning fundamentals<\/li><li>Ability to handle large-scale data<\/li><li>Interest in biomedical applications<\/li><\/ul> \t\t<\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Jonathan Dan, Christodoulos Kechris, Prof. David Atienza<br> Contact email: <a href='mailto:jonathan.dan@epfl.ch; david.atienza@epfl.ch?subject=Extraction of Heart-Related Information in Epileptic Patients'>jonathan.dan@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project660minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Jonathan Dan, Christodoulos Kechris, Prof. David Atienza<br>\";<\/script>\n<span id=project660><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Jonathan Dan, Christodoulos Kechris, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project660',project660); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor659><\/a><b><span style='font-size: 20px;'>Physiological Data Collection Platform for WearOS Smartwatches<\/b><br><script>var project659=\"<div class='entry-content'> \t\t\t <p>The increasing availability of smartwatches with advanced sensors  presents unique opportunities for continuous health monitoring. Building  on our lab&rsquo;s experience with epilepsy monitoring applications, this  project aims to develop a generalized framework for collecting and  managing physiological data from WearOS smartwatch devices. While  current commercial solutions often provide only processed data (like  step counts), there is a need in the research community for access to  raw sensor data. This project will create a flexible platform that  allows researchers to configure and collect both raw and processed  physiological signals from commercially available smartwatches.<\/p>    <p>The technical challenges include managing battery life while  collecting high-frequency sensor data, ensuring cross-device  compatibility, and implementing secure and efficient data transmission  protocols. The system must be user-friendly enough for research  participants while providing the detailed configuration options  researchers need.<\/p>    <p>This project will contribute to the research community by providing a  flexible, open-source platform for physiological data collection using  commercially available smartwatches, enabling various future research  applications beyond epilepsy monitoring.<\/p>    <p>Tasks:<\/p>    <ul class='wp-block-list'><li>Extend the existing smartwatch lab application to support additional sensor data<\/li><li>Implement a configuration interface in the phone companion app for sensor selection and sampling rates<\/li><li>Develop a robust data storage and transmission architecture<\/li><li>Create interfaces for multiple cloud storage services (minimum: Google Drive, Amazon AWS S3)<\/li><li>Test and validate the system on multiple WearOS devices (minimum 3)<\/li><li>Document the system architecture and API<\/li><li>Bonus: Implement real-time data visualization<\/li><li>Bonus: Create or integrate with a researcher dashboard for monitoring data collection<\/li><\/ul>    <p>Requirements:<\/p>    <ul class='wp-block-list'><li>Strong Android development skills (Kotlin)<\/li><li>Understanding of mobile sensor APIs<\/li><li>Knowledge of cloud services and REST APIs<\/li><li>Knowledge of database systems<\/li><li>Experience with Bluetooth communication protocols<\/li><li>Basic understanding of physiological signals<\/li><li>UI\/UX design skills<\/li><\/ul> \t\t<\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Jonathan Dan, Dimitra Tatli, Prof. David Atienza<br> Contact email: <a href='mailto:jonathan.dan@epfl.ch; david.atienza@epfl.ch?subject=Physiological Data Collection Platform for WearOS Smartwatches'>jonathan.dan@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project659minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Jonathan Dan, Dimitra Tatli, Prof. David Atienza<br>\";<\/script>\n<span id=project659><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Jonathan Dan, Dimitra Tatli, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project659',project659); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor636><\/a><b><span style='font-size: 20px;'>Porting X-HEEP to a Lattice FPGA with open-source EDA tools<\/b><br><script>var project636=\"<p>Micro-controller units (MCUs) are used in a wide range of applications ranging from sensor monitoring all the way to robotics. Despite typically lower in performance, they are usually preferred over custom circuits thanks to their versatility and easy programmability via software routines typically written in the C language.<\/p> <p>During the design stage, MCUs are typically implemented in FPGAs to speed up the emulation runtime and allow SW developers to work on the software development kit in parallel with the hardware team.<\/p> <p>Usually, FPGAs are expensive and rely on commercial EDA tools for the synthesis and place-and-route. Lately, however, open-source EDA tools such as YosysHQ have been developed to allow for tool customizations and free costs.<\/p> <p>One of the most supported FPGA vendors supported by YosysHQ is Lattice, which provides very cheap FPGA that can be used for education, rapid prototypes, and all the way to products.<\/p>  <p>In this project, we want to implement in Lattice FPGA using YosysHQ X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform), an open-source, configurable, and extensible single-core RISC-V 32-bit MCU developed at the Embedded Systems Laboratory (ESL), sponsored by the EcoCloud Sustainable Computing center of Swiss Federal Institute of Technology Lausanne (EPFL).<\/p>  <p>In this project, the student will extend the X-HEEP configurations knobs to allow the X-HEEP RTL to be mapped into the Lattice FPGA using the open-source<\/p> <p>synthesis flow based on the YosysHQ EDA tools. This will require several steps to make the RTL HDL &ldquo;digested&rdquo; by the YosysHQ, which supports less advanced HDL features than commercial EDA tools such as Vivado.&nbsp;<\/p>  <p>In particular, the student will need to:<\/p>  <ul> <li>Translate the SystemVerilog description of X-HEEP from SystemVerilog to Verilog with tools such as sv2v<\/li> <li>Add ifdef\/including options to translate only a subset of the supported SystemVerilog RTL to Verilog by sv2v YosysHQ<\/li> <li>Design a synthesis and place-and-route script for YosysHQ to have a working bitstream for the Lattice FPGA<\/li> <li>Load the bitstream into the ICESugar-Pro board that hosts the Lattice FPGA<\/li> <li>Write a simple &ldquo;hello world&rdquo; application for X-HEEP and run it into the ICESugar-Pro board<\/li> <li>[optional] Integrate\/design the SDRAM IP into the X-HEEP RTL so that programs can access the SDRAM present in the ICESugar-Pro board<\/li> <\/ul>  <p>Throughout the project, the student will learn:<\/p>  <ul> <li>How X-HEEP build flow is organized<\/li> <li>How to build a YosysHQ flow for Lattice FPGA for X-HEEP<\/li> <li>How to analyze the output of the open source EDA tools and fix potential violations.<\/li> <li>How to work with several git repositories and in a team of people all contributing to the same project&nbsp;<\/li> <li>How to debug FPGA designs<\/li> <\/ul>  <p>The project will be carried out between the ESL and the TCL groups at EPFL, one of the world&rsquo;s top-class universities. ESL and TCL are active groups (24 Ph.D. students among 45 members for ESL, and 4 Ph.D students among 12 members for TCL) involved in many research aspects. The student will be under the supervision of Prof. David Atienza and Prof. Andreas Burg, Dr. Davide Schiavone, and Dr. Christoph Mueller.<\/p>  <p><strong>Project objectives:<\/strong><\/p> <ol> <li>Understanding the X-HEEP microcontroller, how its build flow works, and learning how IPs are integrated.<\/li> <li>Understanding the sv2v flow to translate SystemVerilog to Verilog.<\/li> <li>How to synthesize, place, and route the translated Verilog with YosysHQ for Lattice FPGAs.<\/li> <li>How to debug and program the design into the ICESugar-Pro board.<\/li> <li>Contribute to the X-HEEP GitHub repository the whole flow.<\/li> <\/ol>  <p><strong>Required knowledge and skills:<\/strong><\/p>  <ul> <li>How the RTL to FPGA flow is usually implemented with the standard flows and tools (e.g. Vivado)<\/li> <li>Python, bash, Linux<\/li> <li>Good analytical skills<\/li> <li>Good background in computer architecture<\/li> <li>Advanced problem-solving skills<\/li> <li>Very curious<\/li> <li>Teamwork and git<\/li> <\/ul>  <p><strong>Appreciated skills:<\/strong><\/p> <ul> <li>Scientific curiosity<\/li> <li>Good communication skills<\/li> <li>Advanced English&nbsp;<\/li> <\/ul> <p><br \/><strong>Type of work:<\/strong> 10% theory analysis, 90% design and simulation<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Christoph Mueller, Prof. David Atienza, Prof. Andreas Burg<br> Contact email: <a href='mailto:davide.schiavone@epfl.ch;christoph.mueller@epfl.ch;david.atienza@epfl.ch;andreas.burg@epfl.ch?subject=Porting X-HEEP to a Lattice FPGA with open-source EDA tools'>davide.schiavone@epfl.ch;christoph.mueller@epfl.ch;david.atienza@epfl.ch;andreas.burg@epfl.ch<\/a><br>\";<\/script>\n<script>var project636minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Christoph Mueller, Prof. David Atienza, Prof. Andreas Burg<br>\";<\/script>\n<span id=project636><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Davide Schiavone, Dr. Christoph Mueller, Prof. David Atienza, Prof. Andreas Burg<br> <a href=#_ onclick=opendesc('project636',project636); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor635><\/a><b><span style='font-size: 20px;'>Testing and verification of a fabricated IMC memory<\/b><br><script>var project635=\"<div class='entry-content'> \t\t\t <p>Efficiently computing complex AI-based workloads on edge devices is a  challenge that industrial and academic groups try to overcome with  various techniques. Among them, computing part of the algorithms in  memory subsystems appears as a competitive solution. This is called In  and Near Memory Computing (IMC, NMC). Focusing on IMC, one of the main  limiting factors resides in the complexity of the design process, as the  subarray itself must be modified. A completely custom memory subarray  has been handcrafted in the Heepocrates chip, designed in 65nm CMOS  technology. This is a tedious and prone-to-error process.<\/p>    <p>This project proposes to characterize the memory and, if this can be  done quickly, contribute to the development of an IMC memory compiler.<\/p>    <p>The memory characterization will be done using a RISC-V core to  control the memory test structure. First, we will characterize the test  structure to ensure its full functionality. Then, test the memory: (i)  perform some functional test to verify its functionality. (ii) write  code to run performance and power tests on the memory. (iii) generate  shmoo plots. Do the same test with several chips to do some statistical  analysis.<\/p>    <p>Then, depending on the available time, we propose to explore the  automation of the physical design of a similar IMC array in 65nm. Python  scripts shall automatically generate a memory subarray in a gdsii file.<\/p>    <p>The project will be carried out at the <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded Systems Laboratory (ESL)<\/a>,  inside the Swiss Federal State Institute of Technology (EPFL), one of  the world&rsquo;s top-class universities. ESL is an active group that is  involved in many research aspects. The student will be under the  supervision of Dr. Alexandre Levisse and Prof. David Atienza.<\/p>    <p><strong>Project objectives:<\/strong><\/p>    <ol class='wp-block-list'><li>Understanding of the BLADE subarray architecture, schematic, and  floorplan. Specifically, it is important to explicitly understand the  operations that can be done in the memory. Understanding of the  characterization structure.<\/li><li>Getting used to the utilization of the test structure developed to  test the memory. Write some basic code and verify the functionality of  the test structure.<\/li><li>Running tests on the memory : (i) test the functionality. (ii) test  the speed. (iii) characterize the power for each operation. (iv) extract  the max freq at various voltages. (v) check several chips. <ol class='wp-block-list'><li>This is mandatory to pass (grade 4). The quality of the generated  code and the efficacy of the student, put in context with the challenges  faced during the project, could bias positively the grade.<\/li><\/ol> <\/li><li>Then, if time allows, we will look into the automation of the array  generation with Python scripts. i.e., generating an arbitrary size SRAM  array with a Python script using the Nazca library.<\/li><\/ol>    <p><strong>Required knowledge and skills:<\/strong><\/p>    <ul class='wp-block-list'><li>C-code and Python.<\/li><li>Linux environment<\/li><li>Good understanding of memory architectures<\/li><li>Advanced knowledge of digital and analog circuit design<\/li><li>Good analytical skills<\/li><li>Good background in computer architecture<\/li><\/ul>    <p><strong>Appreciated skills:<\/strong><\/p>    <ul class='wp-block-list'><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English<\/li><li>Autonomous workability<\/li><li>Teamwork<\/li><\/ul>    <p><strong><u>Type of work:<\/u><\/strong> &nbsp;&nbsp;&nbsp;&nbsp; 20% theory analysis, 80% design and simulation<\/p> \t\t<\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Alexandre Levisse, Prof. David Atienza<br> Contact email: <a href='mailto:alexandre.levisse@epfl.ch;david.atienza@epfl.ch?subject=Testing and verification of a fabricated IMC memory'>alexandre.levisse@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project635minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Alexandre Levisse, Prof. David Atienza<br>\";<\/script>\n<span id=project635><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Alexandre Levisse, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project635',project635); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor632><\/a><b><span style='font-size: 20px;'>Deployment of probabilistic deep learning methods for context-aware and robust Heart Rate extraction in constrained wearables<\/b><br><script>var project632=\"<div class='entry-content'> \t\t\t\t\t <p>Photoplethysmography (PPG) has emerged as a cost-effective  alternative to electrocardiography (ECG) and is now prevalent in a  variety of commercially available wearable devices. This presents a  practical solution for long-term patient monitoring. However, PPG  signals are more susceptible to interference from body motion than ECG,  potentially distorting crucial information about heart functionality,  such as the blood volume pulse (BVP) morphology. To address these motion  artifacts (MA), a plethora of methods have been proposed, with a  predominant focus on accurately extracting heart rate (HR) from PPG. A  proposed solution involves leveraging Deep Learning to integrate both  steps into a single model. This model is designed to simultaneously  perform direct heart rate inference and minimize the effects of motion  artifacts. While these methods aim to diminish the impact of Motion  Artefacts on the ultimate heart rate estimations, it&rsquo;s noteworthy that  no explicit source separation task is outlined during training. Instead,  the implication is that by minimizing the HR-inference loss, the  solution achieved will inherently extract heart rate solely from the  disentangled heart component. The latter fact is false.<\/p>    <p>On this basis, the goal of this project is to employ a novel  probabilistic-driven deep learning methodology, which has been designed  and developed at the Embedded Systems Laboratory (ESL) at EPFL.  Specifically, a design space exploration will be performed, given the  restrictions of the application and the HW\/SW capabilities of the  available platform. This will result in a clear understanding of the  bottlenecks and limitations of the proposed deep learning architecture  when trying to deploy them in extreme edge ultra-low-power devices.<\/p>    <p>Throughout the project, the student will learn:<\/p>    <ul class='wp-block-list'><li>How to interact with cutting-edge System-on-Chips.<\/li><li>How to deal with embedded systems (memories, peripherals, etc).<\/li><li>How to write optimal C-level.<\/li><li>How to design and develop efficient embedded APIs.<\/li><li>How to properly debug embedded applications.<\/li><li>How to work with git repositories.<\/li><li>How to interface with the other people of the team (machine learning, etc.) contributing to the project.<\/li><\/ul>    <p>The project will be carried out at the ESL at EPFL, one of the  world&rsquo;s top-class universities including EcoCloud&rsquo;s technical support.  ESL is an active group (24 Ph.D. students among 45 members) involved in  many research aspects. The student will be under the supervision of Christodoulos Kechris MSc., Dr. Jos&eacute; Miranda, and Dr. Jonathan Dan, as key  senior daily supervisors,&nbsp; and Prof. David Atienza.<\/p>    <p><strong>Project objectives:<\/strong><\/p>    <ol class='wp-block-list'><li>Understanding the current deep learning model(s) to be deployed, and  proposing a step-by-step\/white-box planning to reach its deployment  into an embedded constrained platform.<\/li><li>Developing and integrating the different stages of the probabilistic deep learning pipeline.<\/li><li>Validation of previous point given specific public available data.<\/li><li>Analysis of both memory and inference timings given the performed implementation.<\/li><li>[Optional]. Utilization and analysis of different HW\/SW optimisations to improve the metrics of the previous point.<\/li><li>[Optional]. Plugging-in the already implemented first stages of the architecture and validation of the whole system.<\/li><li>[Optional]. Testing the architecture using a real Body Area Network system available at the lab (e.g.: VersaSens).<\/li><\/ol>    <p><strong>Required knowledge and skills:<\/strong><\/p>    <ul class='wp-block-list'><li>Low-level software design (C and\/or C++ is going to be used throughout the project)<\/li><li>Good understanding of memory architectures and microcontrollers<\/li><li>Good analytical skills<\/li><li>Good background in computer architecture and algorithms<\/li><li>Teamwork and git<\/li><\/ul>    <p><strong>Appreciated skills:<\/strong><\/p>    <ul class='wp-block-list'><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English<\/li><\/ul>    <p><strong>Type of work:<\/strong> 10% theory analysis, 90% design and simulation<\/p> \t\t\t\t\t\t\t\t\t<\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Christodoulos Kechris, Dr. Jose Miranda, Dr. Jonathan Dan, Prof. David Atienza <br> Contact email: <a href='mailto:christodoulos.kechris@epfl.ch; jose.mirandacalero@epfl.ch; jonathan.dan@epfl.ch; david.atienza@epfl.ch?subject=Deployment of probabilistic deep learning methods for context-aware and robust Heart Rate extraction in constrained wearables'>christodoulos.kechris@epfl.ch; jose.mirandacalero@epfl.ch; jonathan.dan@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project632minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Christodoulos Kechris, Dr. Jose Miranda, Dr. Jonathan Dan, Prof. David Atienza <br>\";<\/script>\n<span id=project632><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Christodoulos Kechris, Dr. Jose Miranda, Dr. Jonathan Dan, Prof. David Atienza <br> <a href=#_ onclick=opendesc('project632',project632); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor594><\/a><b><span style='font-size: 20px;'>Towards Audio-Visual Speech Separation and Recognition<\/b><br><script>var project594=\"<div class='entry-content'>   <p>Human-computer interaction evolved rapidly in the last years thanks  to the hardware development enabling complex Deep Neural Networks (DNNs)  to perform machine learning tasks. The scientific advancements led to  the employment of Automated Speech Recognition (ASR) and Visual Speech  Recognition (VSR) in human-machine communication, with the former  representing the task of converting spoken speech into written words  [2], whilst the latter being the conversion of lip movement into written  words (i.e., lipreading) [6]. In the context of voice-controlled  devices, it is of uttermost importance to additionally mention speaker  separation [7]. This aim to mimic the cocktail party effect present in  humans, which denotes our capability of focusing on an individual  conversation, whilst filtering out other discussions and surrounding  noises. This ensures that only a designated individual can control one  target device at a time. On this basis, the <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded System Laboratory<\/a>  at Ecole Polytechnique F&eacute;d&eacute;rale de Lausanne (EPFL) together with the  Integrated System Laboratory at Eidgen&ouml;ssische Technische Hochschule  Z&uuml;rich (ETH) are developing different platforms, tools, and frameworks  to tackle the next-generation IoT audio devices. In the current work, we  investigate the potential of performing Audio-Visual Speech Separation  and Recognition (AVSSR) using sensor fusion [3, 4, 8]. Through  multi-task learning, we merge Audio-Visual Speech Separation (AVSS) and  Audio-Visual Speech Recognition (AVSR). We propose to fuse audio and  visual information, thus increasing the data dimensionality and,  implicitly, the amount of useful information. The so-trained model would  then provide the target user(s) with the transcript of their respective  speech. Lastly, the proposed system must abide by the TinyML [5]  constraints considering edge devices, namely reduced memory and storage  requirements, as well as real-time operation on low-power,  battery-operated devices.<\/p>             <p>The project will be carried out at the ESL at EPFL, one of the  world&rsquo;s top-class universities. ESL is an active group (24 Ph.D.  students among 45 members) involved in many research aspects. The  student will be under the supervision of Dr. Jos&eacute; Miranda, Mr. Cristian  Cioflan, and Dr. Miguel de Prado, as key senior daily supervisors,&nbsp; and  Prof. David Atienza.<\/p>    <p><strong>Project objectives:<\/strong><\/p>    <ol><li>Familiarize yourself with the project specifics (1-2 Weeks) <ol style='list-style-type: lower-alpha'><li>Learn about DNN training and PyTorch, how to visualize results with TensorBoard.<\/li><li>Read up on data fusion and multimodal learning, common approaches and recent advances on the topic.<\/li><li>Read up on multi-task learning in the context of audio-visual time series.<\/li><li>Read up on DNN models aimed at time series (e.g., TCNs, TASMs,  Transformer and Conformer networks) and the recent advances in  AVSR\/AVSS.<\/li><\/ol> <\/li><li>Propose and evaluate AVSSR topologies (4-6 weeks) <ol style='list-style-type: lower-alpha'><li>Considering state-of-the-art works on AVSS[9] and AVSR[10][11][12]  and previous IIS projects, propose and implement AVSSR architectures.<\/li><li>Propose evaluation metrics and novel loss functions; analyse the models&rsquo; performance on the GRID dataset [1]<\/li><\/ol> <\/li><li>Optimize proposed models considering TinyML constraints (2-3 weeks) <ol style='list-style-type: lower-alpha'><li>Reduce the models&rsquo; hardware-associated costs (i.e., memory, storage,  computational complexity), evaluating the trade-offs on the proposed  metrics.<\/li><li>(Only if conducted as a Master&rsquo;s thesis) Deploy the proposed architecture on novel ultra-low-power platforms.<\/li><\/ol> <\/li><li>(Optional) Dataset generalization and ablation study (1-2 Weeks) <ol style='list-style-type: lower-alpha'><li>Investigate alternative datasets for AVSSR.<\/li><li>Propose and implement dataset modifications to enable multi-speaker AVSSR.<\/li><li>Evaluate and compare multimodal learning against audio- and video-only learning.<\/li><li>Evaluate and compare speaker-overlapping and speaker-disjoint training and testing.<\/li><\/ol> <\/li><li>Gather and Present Final Results (2-3 Weeks): Write a final report.  Include all major decisions taken during the design process and argue  your choice. Include everything that deviates from the very standard  case, show off everything that took time to figure out and all your  ideas that have influenced the project.<\/li><\/ol>    <p><strong>Required knowledge and skills:<\/strong><\/p>    <ul><li>Python programming<\/li><li>Deep learning knowledge<\/li><li>Automatic speech recognition understanding\/experience<\/li><li>Strong background in computer science<\/li><li>Teamwork and git<\/li><\/ul>    <p><strong>Appreciated skills:<\/strong><\/p>    <ul><li>Scientific curiosity, good communication skills, and advanced English<\/li><\/ul>    <p><strong>Type of work:<\/strong> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 40% research and theory analysis, 60% development and implementation<\/p>    <p><strong>References<\/strong><\/p>    <p>[1]&nbsp;Mishaim Malik, Muhammad Kamran Malik, Khawar Mehmood, and Imran Makhdoom&nbsp;Automatic speech recognition: a survey.&nbsp;2021.<\/p>    <p>[2]&nbsp;Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng  Yu, Dong Yu, and Jesper Jensen.&nbsp;An overview of deep-learning-based  audio-visual speech enhancement and separation.&nbsp;2020.<\/p>    <p>[3]&nbsp;Liliane Momeni, Triantafyllos Afouras, Themos Stafylakis, Samuel  Albanie, and Andrew Zisserman.&nbsp;Seeing wake words: Audio-visual keyword  spotting.&nbsp;2020.<\/p>    <p>[4]&nbsp;Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev  Khudanpu.&nbsp;Librispeech: An asr corpus based on public domain audio  books.&nbsp;2015.<\/p>    <p>[5]&nbsp;Vijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson,  Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien  Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam  Davis, Pan Deng, Greg Diamos, Jared Duke, Dave Fick, J. Scott Gardner,  Itay Hubara, Sachin Idgunji, Thomas B. Jablin, Jeff Jiao, Tom St. John,  Pankaj Kanwar, David Lee, Jeffery Liao, Anton Lokhmotov, Francisco  Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady  Pekhimenko, Arun Tejusve Raghunathm Rajan, Dilip Sequeira, Ashish  Sirasao, Fei Sun, Hanlin Tang, Michael Thomson, Frank Wei, Ephrem Wu,  Lingjie Xu, Koichi Yamada, Bing Yu, George Yuan, Aaron Zhong, Peizhao  Zhang, and Yuchen Zhou.&nbsp;&ldquo;MLperf inference benchmark.&nbsp;2020<\/p>    <p>[6]&nbsp;Changchong Sheng, Gangyao Kuang, Liang Bai, Chenping Hou, Yulan  Guo, Xin Xu, Matti Pietik&auml;inen, and Li Liu.&nbsp;Deep learning for visual  speech analysis: A survey,&nbsp;2022<\/p>    <p>[7]&nbsp;DeLiang Wang and Jitong Chen.&nbsp;Supervised speech separation based on deep learning: An overview.&nbsp;2018<\/p>         <\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Cristian Cioflan, Dr. Miguel de Prado, Dr. Jose Miranda, Prof. David Atienza<br> Contact email: <a href='mailto:cioflanc@iis.ee.ethz.ch; miguel.deprado@verses.ai; jose.mirandacalero@epfl.ch; david.atienza@epfl.ch?subject=Towards Audio-Visual Speech Separation and Recognition'>cioflanc@iis.ee.ethz.ch; miguel.deprado@verses.ai; jose.mirandacalero@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project594minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Cristian Cioflan, Dr. Miguel de Prado, Dr. Jose Miranda, Prof. David Atienza<br>\";<\/script>\n<span id=project594><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Cristian Cioflan, Dr. Miguel de Prado, Dr. Jose Miranda, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project594',project594); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor591><\/a><b><span style='font-size: 20px;'>Implementation of automation techniques for IMC arrays<\/b><br><script>var project591=\"<div class='entry-content mb-5'> \t\t <p>Efficiently computing complex AI-based workloads on edge devices is a  challenge that industrial and academic groups try to overcome with  various techniques. Among them, computing part of the algorithms in  memory subsystems appears as a competitive solution. This is called In  and Near Memory Computing (IMC, NMC). Focusing on IMC, one of the main  limiting factors resides in the complexity of the design process, as the  subarray itself must be modified. In the Heepocrates chip, designed in  65nm CMOS technology, a completely custom memory subarray has been  handcrafted. This is a tedious and prone-to-error process.<\/p> <p>This project proposes to lay the foundations of an IMC memory  compiler by exploring the automation of some parts of the subarray,  specifically the decoder and the controller. Starting from an HDL  definition of these blocks, the objective of the project is to  parametrize them, automate their physical design, and integrate them  inside the subarray. Finally, if time allows, exploring the utilization  of open-source tool flows could be considered.<\/p> <p>The project will be carried out at the <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded Systems Laboratory (ESL)<\/a>,  inside the Swiss Federal State Institute of Technology (EPFL), one of  the world&rsquo;s top-class universities. ESL is an active group involved in  many research aspects. The student will be under the supervision of  Prof. David Atienza and Dr. Alexandre Levisse.<\/p>  <p><strong>Project objectives:<\/strong><\/p> <ol><li>Understanding of the BLADE subarray architecture, schematic, and  floorplan. Specifically understanding the details of the decoder and  subarray controller. Definition of the top entities (input, outputs,  functional behavior, metrics).<\/li><li>Design of the RTL code of the decoder. Synthesis, PnR under physical  constraints to fit it on the available area enclosure and pitch  matching of the outputs with the array drivers. Verification with spice  simulations.<\/li><li>Utilization of the same flow on the subarray controller. Definition  of timing constraints. Floorplan updates. Update of the array schematic  and layout.<\/li><li>If time allows, the project targets the utilization of open source synthesis and PnR tools in the flow.<\/li><\/ol>  <p><strong>Required knowledge and skills:<\/strong><\/p> <ul><li>Good understanding of memory architectures<\/li><li>Advanced knowledge of digital and analog circuit design<\/li><li>Good analytical skills<\/li><li>Good background in computer architecture<\/li><\/ul>  <p><strong>Appreciated skills:<\/strong><\/p> <ul><li>Scientific curiosity<\/li><li>Good communication skills<\/li><li>Advanced English<\/li><li>Autonomous workability<\/li><li>Teamwork<\/li><\/ul>  <p><strong><u>Type of work:<\/u><\/strong> &nbsp;&nbsp;&nbsp;&nbsp; 20% theory analysis, 80% design and simulation<\/p> \t<\/div><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Alexandre Levisse, Prof. David Atienza <br> Contact email: <a href='mailto:alexandre.levisse@epfl.ch; david.atienza@epfl.ch?subject=Implementation of automation techniques for IMC arrays'>alexandre.levisse@epfl.ch; david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project591minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Alexandre Levisse, Prof. David Atienza <br>\";<\/script>\n<span id=project591><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Dr. Alexandre Levisse, Prof. David Atienza <br> <a href=#_ onclick=opendesc('project591',project591); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor589><\/a><b><span style='font-size: 20px;'>Deployment of an unsupervised deep adaptive filtering method for context-aware and robust Heart Rate extraction in constrained wearables<\/b><br><script>var project589=\"<p>Photoplethysmography (PPG) has emerged as a cost-effective alternative to electrocardiography (ECG) and is now prevalent in a variety of commercially available wearable devices. This presents a practical solution for long-term patient monitoring. However, PPG signals are more susceptible to interference from body motion than ECG, potentially distorting crucial information about heart functionality, such as the blood volume pulse (BVP) morphology. To address these motion artifacts (MA), a plethora of methods have been proposed, with a predominant focus on accurately extracting heart rate (HR) from PPG. A proposed solution involves leveraging Deep Learning to integrate both steps into a single model. This model is designed to simultaneously perform direct heart rate inference and minimize the effects of motion artifacts. While these methods aim to diminish the impact of Motion Artefacts on the ultimate heart rate estimations, it\u2019s noteworthy that no explicit source separation task is outlined during training. Instead, the implication is that by minimizing the HR-inference loss, the solution achieved will inherently extract heart rate solely from the disentangled heart component. The latter fact is false.&nbsp;<\/p> <p>On this basis, the goal of this project is to employ unsupervised deep adaptive filtering methods, which has been designed and developed at the Embedded Systems Laboratory (ESL) at EPFL, deploying it into heterogeneous SoCs based on RISC-V ISA. For instance, into X-HEEP, (eXtendable Heterogeneous Energy-Efficient Platform), which is an open-source, configurable, and extensible single-core RISC-V 32b MCU, sponsored by the EcoCloud Sustainable Computing center of EPFL. It is based on many third-party open-source IPs and in-house IPs developed at ESL jointly with other EPFL laboratories. X-HEEP provides a framework to run applications compiled for RISC-V on a simulator (Verilator, Questasim, or VCS), on a Xilinx FPGA, and can be implemented in silicon as well. Specifically, a design space exploration will be performed, given the restrictions of the application and the HW\/SW capabilities of the platform. This will result in a clear understanding of the bottlenecks and limitations of the proposed deep learning architecture when trying to deploy them in extreme edge ultra-low-power devices.<\/p> <p>Throughout the project, the student will learn:<\/p> <ul> <li>How to interact with cutting-edge System-on-Chips.<\/li> <li>How to deal with embedded systems (memories, peripherals, etc).<\/li> <li>How to write optimal C-level.<\/li> <li>How to design and develop efficient embedded APIs.<\/li> <li>How to properly debug embedded applications.<\/li> <li>How to work with git repositories.<\/li> <li>How to interface with the other people of the team (machine learning, etc.) contributing to the project.<\/li> <\/ul> <p>The project will be carried out at the ESL at EPFL, one of the world\u2019s top-class universities including EcoCloud\u2019s technical support. ESL is an active group (24 Ph.D. students among 45 members) involved in many research aspects. The student will be under the supervision of MSc. Christodoulos Kechris, Dr. Jos\u00e9 Miranda, and Dr. Jonathan Dan, as key senior daily supervisors,&nbsp; and Prof. David Atienza.&nbsp;<\/p> <p><strong>Project objectives:<\/strong><\/p> <ol> <li>Understanding the current unsupervised adaptive deep learning model to be deployed.<\/li> <li>Developing and integrating the initial stages of the deep learning architecture into X-HEEP.<\/li> <li>Validation of previous point given specific public available data.<\/li> <li>Developing and integrating the latter stages of the architecture.<\/li> <li>Validation of the whole system.<\/li> <li>Analysis of both memory and training\/inference timings given the performed implementation.&nbsp;<\/li> <li>[Optional]. Utilization and analysis of different HW\/SW optimisations to improve the metrics of the previous point.<\/li> <\/ol> <p><strong>Required knowledge and skills:<\/strong><\/p> <ul> <li>Low-level software design (C and\/or C++ is going to be used throughout the project)<\/li> <li>Good understanding of memory architectures and microcontrollers<\/li> <li>Good analytical skills<\/li> <li>Good background in computer architecture and algorithms<\/li> <li>Teamwork and git<\/li> <\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul> <li>Scientific curiosity<\/li> <li>Good communication skills<\/li> <li>Advanced English&nbsp;<\/li> <\/ul> <p><strong>Type of work:<\/strong> 5% theory analysis, 95% design and simulation<\/p><br><b>Lab: <\/b>ESL <br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Christodoulos Kechris MSc, Dr. Jose Miranda, Dr. Jonathan Dan, Prof. David Atienza<br> Contact email: <a href='mailto:christodoulos.kechris@epfl.ch;jose.mirandacalero@epfl.ch;jonathan.dan@epfl.ch;david.atienza@epfl.ch?subject=Deployment of an unsupervised deep adaptive filtering method for context-aware and robust Heart Rate extraction in constrained wearables'>christodoulos.kechris@epfl.ch;jose.mirandacalero@epfl.ch;jonathan.dan@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project589minus=\"<b>Lab: <\/b>ESL <br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Christodoulos Kechris MSc, Dr. Jose Miranda, Dr. Jonathan Dan, Prof. David Atienza<br>\";<\/script>\n<span id=project589><b>Lab: <\/b>ESL <br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Christodoulos Kechris MSc, Dr. Jose Miranda, Dr. Jonathan Dan, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project589',project589); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><a href=# title='theoretical project'><img width=30 alt='theoretical' src=https:\/\/eslweb.epfl.ch\/projects\/images\/t.gif hspace=2 border=0><\/a><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><td width=10 rowspan=2 valign=top><\/td><td rowspan=2 valign=top width=600 style='min-width:600px !important;'><a style='position: relative; top:-30px' name=anchor588><\/a><b><span style='font-size: 20px;'>Design of efficient Convolutional Layers for Near-Memory Computing IPs based on RISC-V<\/b><br><script>var project588=\"<p><em>Artificial Intelligence<\/em> (AI) has been one of the most dominant factors driving technology innovation over the last decade. Cars that drive autonomously, cameras that recognize faces, microphones that recognize voice commands, and wearable health monitors that detect epilepsy attacks are only a few examples born from the AI revolution. As silicon devices become smaller and faster, system-on-chips (SoCs) become more and more complex, enabling pocket-size, wearable, battery-powered systems to compute millions\/billions of operations per second in a limited power budget. To support a large variety of rapidly changing <em>AI<\/em>-backed applications, software programmable SoCs are usually preferred to more specialized architectures thanks to their versatility and short time-to-market.<\/p> <p>One of the main nowadays SoCs performance and energy limitations resides in the high memory traffic generated by traditional Von Neuman systems, which is one of the critical bottlenecks to solve for the next generation of smart systems.<\/p> <p>One promising idea to overcome this limitation is to bring computation into the memory subsystem to use the available bandwidth more efficiently and leverage data reutilization.<br>Such a computational paradigm is known as In-Memory (or Near-Memory) computing.<\/p> <p>The <a href='https:\/\/www.epfl.ch\/labs\/esl\/'>Embedded Systems Laboratory (ESL)<\/a> at Swiss Federal Institute of Technology Lausanne (EPFL) has developed a TSMC 65nm low-power SRAM-based programmable near-computing architecture (known as Carus) capable of computing operations (such as addition, and, or, xor, multiplication, etc.) between two words within the memory layout without moving the operands out into processing elements outside the memory. This reduces the number of bus transactions, resulting in a lower energy consumption when compared to Von Neumann architectures such as CPUs.<\/p> <p>Carus works in two different modes: memory mode, where it behaves exactly like a normal memory; and computing mode, where Carus executes previously loaded programs, typically written in RISC-V assembly.<\/p> <p>Carus has been integrated into a 32b microcontroller called X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform). X-HEEP is an open-source, configurable, and extensible single-core RISC-V 32b MCU, sponsored by the EcoCloud Sustainable Computing center of EPFL. It is based on many third-party open-source IPs and in-house IPs developed at the Embedded Systems Laboratory (ESL) jointly with other EPFL laboratories.<\/p> <p>X-HEEP provides a framework to run applications compiled for RISC-V on a simulator (Verilator, Questasim, or VCS), on a Xilinx FPGA, and can be implemented in silicon as well.<\/p> <p>To execute kernels, the system RISC-V CPU sets Carus into computing mode, loads the appropriate Carus program, copies the input data to Carus\u2019s memory (leveraging the DMA), and finally triggers Carus execution.<\/p> <p>As AI is dominating edge-computing workloads, Carus has been designed to run efficiently convolutional layers. However, due to the number of parameters, such as the size of the inputs and weights filters, number of channels, filters, padding, strides, data types, etc, optimally partitioning and mapping such layers is not trivial, and requires an in-depth study and characterization of the impact of the available data-flow variants (input, output, or weight stationary), tiling strategies, etc.<\/p> <p>Therefore, we propose to develop a library of computing kernels that support different parameters and data flows to efficiently deploy convolutional layers on the system by leveraging Carus computing capabilities.  Such library takes as input the activation and filter dimensions and other parameters such as stride, padding, data-flow, etc. Then, the system CPU applies an appropriate tiling to the tensors to maximise the exploitation of Carus available memory and computing bandwidth. The goal is to profile Carus\u2019s performance under different parameters and build a performance and energy model of the IP when executing convolutional layers.<\/p> <p>Throughout the project, the student will learn:<\/p> <ul> <li>how to partition and optimize Convolutional Layer software implementations using the system CPU.<\/li> <li>how to design and optimize Convolutional Layer with the Carus IP.<\/li> <li>how to analyze, profile, and improve the performance of the designed library.<\/li> <li>how to model a system to estimate its performance given a set of input parameters and constraints.<\/li> <li>how to work with version control (Git) and third-party, open-source repositories and tools.<\/li> <li>How to work in a team of people all contributing to the same project.<\/li> <\/ul> <p>The project will be carried out at the ESL at EPFL, one of the world\u2019s top-class universities. ESL is an active group (24 Ph.D. students among 45 members) involved in many research aspects, therefore providing a stimulating research environment. The student will be under the supervision of Prof. David Atienza, Dr. Davide Schiavone, and Mr. Luigi Giuffrida.<\/p> <p><strong>Project objectives:<\/strong><\/p> <p>Project objectives:<\/p> <ol> <li>Understand the architecture and working principles of the X-HEEP microcontroller, and learn how IPs are integrated into it.<\/li> <li>Understand the Carus memory architecture and learn how it is connected and leveraged in the X-HEEP\/HEEPerator MCU.<\/li> <li>Develop a C implementation of a tiled Convolutional Layer using the system C, supporting different data flows and leveraging the existing kernel implementations to offload computation to Carus. Analyze the performance.<\/li> <li>Develop the Carus\u2019s Assembly implementation of a Convolutional Layer using different data flow, and analyze the performance.<\/li> <li>Validate, Profile, and Analyze the new software library on HEEPerator when using it to run real-world neural networks.<\/li> <\/ol> <p>Throughout the project, the student will learn:<\/p> <ul> <li>How to optimize C code targeting execution speed on near-memory devices.<\/li> <li>How to write assembly code for a custom computing engine based on RISC-V.<\/li> <li>The relevance of hardware-software co-design and its impact on performance.<\/li> <li>Different profiling and characterization techniques.<\/li> <li>How to efficiently analyze and model experimental results.<\/li> <\/ul> <p><strong>Required knowledge and skills:<\/strong><\/p> <ul> <li>Low-level, embedded system C programming.<\/li> <li>Strong analytical skills.<\/li> <li>Strong background in computer architecture and algorithms.<\/li> <li>Teamwork and version control using Git.<\/li> <\/ul> <p><strong>Appreciated skills:<\/strong><\/p> <ul> <li>Scientific curiosity, good communication skills, and advanced English.<\/li> <\/ul> <p><strong>Type of work:<\/strong><\/p> <p>10% research and theory analysis, 60% development and implementation, 30% analysis of results.<\/p><br><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Luigi Giuffrida, Dr. Davide Schiavone, Prof. David Atienza<br> Contact email: <a href='mailto:luigi.giuffrida@epfl.ch;davide.schiavone@epfl.ch;david.atienza@epfl.ch?subject=Design of efficient Convolutional Layers for Near-Memory Computing IPs based on RISC-V'>luigi.giuffrida@epfl.ch;davide.schiavone@epfl.ch;david.atienza@epfl.ch<\/a><br>\";<\/script>\n<script>var project588minus=\"<b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Luigi Giuffrida, Dr. Davide Schiavone, Prof. David Atienza<br>\";<\/script>\n<span id=project588><b>Lab: <\/b>ESL<br><b>Sections: <\/b>SEL<br><b>Supervisor<\/b>: Mr. Luigi Giuffrida, Dr. Davide Schiavone, Prof. David Atienza<br> <a href=#_ onclick=opendesc('project588',project588); style='position:relative;z-index:99;'>[read&nbsp;on]<\/a><\/span><br><hr style='border-top: 1px solid #999; color: #999; background-color: #fff; height: 1px; margin: 20px 0;'><br><\/td><td valign=top width=100><table cellpadding=0 cellspacing=0 border=0><td><a href=# title='experimental project'><img alt='experimental project' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/e.gif hspace=2 border=0><\/a><\/td><td><img width=27 src=https:\/\/eslweb.epfl.ch\/img\/1pixel.gif><\/td><td><a href=# title='computational project'><img alt='computational' width=30 src=https:\/\/eslweb.epfl.ch\/projects\/images\/c.gif hspace=2 border=0><\/a><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/docnolink.gif><\/td><td><img width=30 hspace=2 src=https:\/\/eslweb.epfl.ch\/projects\/images\/webnolink.gif><\/td><\/table><\/td><\/tr><tr><td><a href=https:\/\/www.epfl.ch\/labs\/esl\/research\/systems-on-chip\/x-heep\/ target=_blank title='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><img src=https:\/\/eslweb.epfl.ch\/img\/collaborations\/industry\/201.png width=70 alt='eXtendable Heterogeneous Energy-Efficient Platform - EPFL'><\/a><\/td><\/tr><\/table><\/div><\/div><\/div>\n","protected":false},"excerpt":{"rendered":"<p>Edit Master Projects<\/p>\n","protected":false},"author":7,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-119","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=\/wp\/v2\/pages\/119","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=119"}],"version-history":[{"count":3,"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=\/wp\/v2\/pages\/119\/revisions"}],"predecessor-version":[{"id":138,"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=\/wp\/v2\/pages\/119\/revisions\/138"}],"wp:attachment":[{"href":"https:\/\/eslweb.epfl.ch\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=119"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}