Erik Hagersten
Professor emeritus i datorarkitektur 99 vid Institutionen för informationsteknologi; Datorteknik
- Mobiltelefon:
- 079-321 67 71
- E-post:
- Erik.Hagersten@it.uu.se
- Besöksadress:
- Hus 10, Regementsvägen 10
- Postadress:
- Box 524
751 20 UPPSALA
- Akademiska meriter:
- FD
Publikationer
Senaste publikationer
-
Directed Statistical Warming through Time Traveling
Ingår i MICRO'52, s. 1037-1049, 2019
-
Delorean: Virtualized Directed Profiling for Cache Modeling in Sampled Simulation
2018
-
Tail-PASS: Resource-based Cache Management for Tiled Graphics Rendering Hardware
Ingår i Proc. 16th International Conference on Parallel and Distributed Processing with Applications, s. 55-63, 2018
-
2017
-
A split cache hierarchy for enabling data-oriented optimizations
Ingår i Proc. 23rd International Symposium on High Performance Computer Architecture, s. 133-144, 2017
Alla publikationer
Artiklar i tidskrift
-
Exploring scheduling effects on task performance with TaskInsight
Ingår i Supercomputing frontiers and innovations, s. 91-98, 2017
-
Ingår i IEEE Transactions on Computers, s. 3537-3551, 2016
-
Building Heterogeneous Unified Virtual Memories (UVMs) without the Overhead
Ingår i ACM Transactions on Architecture and Code Optimization (TACO), 2016
- DOI för Building Heterogeneous Unified Virtual Memories (UVMs) without the Overhead
- Ladda ner fulltext (pdf) av Building Heterogeneous Unified Virtual Memories (UVMs) without the Overhead
-
The effects of granularity and adaptivity on private/shared classification for coherence
Ingår i ACM Transactions on Architecture and Code Optimization (TACO), 2015
-
Reconsidering algorithms for iterative solvers in the multicore era
Ingår i International Journal of Computational Science and Engineering, s. 270-282, 2009
-
Fast Data-Locality Profiling of Native Execution
Ingår i ACM SIGMETRICS Performance Evaluation Review, s. 169-180, 2005
-
Parallella program ger paradigmskifte
Ingår i Elektroniktidningen, 2005
Kapitel i böcker, delar av antologi
-
Efficient cache modeling with sparse data
Ingår i Processor and System-on-Chip Simulation, s. 193-209, Springer, 2010
-
TImestamp-based Selective Cache Allocation
Ingår i High Performance Memory Systems, Springer-Verlag, 2003
Konferensbidrag
-
Directed Statistical Warming through Time Traveling
Ingår i MICRO'52, s. 1037-1049, 2019
-
Tail-PASS: Resource-based Cache Management for Tiled Graphics Rendering Hardware
Ingår i Proc. 16th International Conference on Parallel and Distributed Processing with Applications, s. 55-63, 2018
-
A split cache hierarchy for enabling data-oriented optimizations
Ingår i Proc. 23rd International Symposium on High Performance Computer Architecture, s. 133-144, 2017
-
Understanding the interplay between task scheduling, memory and performance
Ingår i Proc. Companion 8th ACM International Conference on Systems, Programming, Languages, and Applications, s. 21-23, 2017
-
A graphics tracing framework for exploring CPU+GPU memory systems
Ingår i Proc. 20th International Symposium on Workload Characterization, s. 54-65, 2017
-
POSTER: Putting the G back into GPU/CPU Systems Research
Ingår i 2017 26TH INTERNATIONAL CONFERENCE ON PARALLEL ARCHITECTURES AND COMPILATION TECHNIQUES (PACT), s. 130-131, 2017
-
Approximation: A New Paradigm also for Wireless Sensing
2016
-
CoolSim: Statistical Techniques to Replace Cache Warming with Efficient, Virtualized Profiling
Ingår i Proceedings Of 2016 International Conference On Embedded Computer Systems, s. 106-115, 2016
-
Formalizing data locality in task parallel applications
Ingår i Algorithms and Architectures for Parallel Processing, s. 43-61, 2016
-
CoolSim: Eliminating Traditional Cache Warming with Fast, Virtualized Profiling
Ingår i 2016 IEEE International Symposium On Performance Analysis Of Systems And Software ISPASS 2016, s. 149-150, 2016
-
Data placement across the cache hierarchy: Minimizing data movement with reuse-aware placement
Ingår i Proc. 34th International Conference on Computer Design, s. 117-124, 2016
-
Effects of Granularity/Adaptivity on Private/Shared Classification for Coherence
2015
-
Long Term Parking (LTP): Criticality-aware Resource Allocation in OOO Processors
Ingår i Proc. 48th International Symposium on Microarchitecture, s. 334-346, 2015
-
StatTask: Reuse distance analysis for task-based applications
Ingår i Proc. 7th Workshop on Rapid Simulation and Performance Evaluation, s. 1-7, 2015
-
Full speed ahead: Detailed architectural simulation at near-native speed
Ingår i Proc. 18th International Symposium on Workload Characterization, s. 183-192, 2015
-
Micro-Architecture Independent Analytical Processor Performance and Power Modeling
Ingår i 2015 IEEE International Symposium on Performance Analysis and Software (ISPASS), s. 32-41, 2015
-
AREP: Adaptive Resource Efficient Prefetching for Maximizing Multicore Performance
Ingår i Proc. 24th International Conference on Parallel Architectures and Compilation Techniques, s. 367-378, 2015
- DOI för AREP: Adaptive Resource Efficient Prefetching for Maximizing Multicore Performance
- Ladda ner fulltext (pdf) av AREP: Adaptive Resource Efficient Prefetching for Maximizing Multicore Performance
-
An efficient, self-contained, on-chip directory: DIR1-SISD
Ingår i Proc. 24th International Conference on Parallel Architectures and Compilation Techniques, s. 317-330, 2015
-
Cost-effective speculative scheduling in high performance processors
Ingår i Proc. 42nd International Symposium on Computer Architecture, s. 247-259, 2015
-
Extending statistical cache models to support detailed pipeline simulators
Ingår i 2014 IEEE International Symposium On Performance Analysis Of Systems And Software (Ispass), s. 86-95, 2014
-
A software based profiling method for obtaining speedup stacks on commodity multi-cores
Ingår i 2014 IEEE INTERNATIONAL SYMPOSIUM ON PERFORMANCE ANALYSIS OF SYSTEMS AND SOFTWARE (ISPASS), s. 148-157, 2014
-
A case for resource efficient prefetching in multicores
Ingår i Proc. International Symposium on Performance Analysis of Systems and Software, s. 137-138, 2014
-
A case for resource efficient prefetching in multicores
Ingår i Proc. 43rd International Conference on Parallel Processing, s. 101-110, 2014
-
The Direct-to-Data (D2D) Cache: Navigating the cache hierarchy with a single lookup
Ingår i Proc. 41st International Symposium on Computer Architecture, s. 133-144, 2014
-
Resource conscious prefetching for irregular applications in multicores
Ingår i Proc. International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation (SAMOS XIV), s. 34-43, 2014
-
The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence
2014
-
Bandwidth Bandit: Quantitative Characterization of Memory Contention
Ingår i Proc. 11th International Symposium on Code Generation and Optimization, s. 99-108, 2013
-
TLC: A tag-less cache for reducing dynamic first level cache energy
Ingår i Proceedings of the 46th International Symposium on Microarchitecture, s. 49-61, 2013
-
Modeling performance variation due to cache sharing
Ingår i Proc. 19th IEEE International Symposium on High Performance Computer Architecture, s. 155-166, 2013
- DOI för Modeling performance variation due to cache sharing
- Ladda ner fulltext (pdf) av Modeling performance variation due to cache sharing
-
Low Overhead Instruction-Cache Modeling Using Instruction Reuse Profiles
Ingår i International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD'12), s. 260-269, 2012
-
Phase Guided Profiling for Fast Cache Modeling
Ingår i International Symposium on Code Generation and Optimization (CGO'12), s. 175-185, 2012
-
Phase Behavior in Serial and Parallel Applications
Ingår i International Symposium on Workload Characterization (IISWC'12), 2012
-
Efficient techniques for predicting cache sharing and throughput
Ingår i Proc. 21st International Conference on Parallel Architectures and Compilation Techniques, s. 305-314, 2012
- DOI för Efficient techniques for predicting cache sharing and throughput
- Ladda ner fulltext (pdf) av Efficient techniques for predicting cache sharing and throughput
-
Bandwidth bandit: Quantitative characterization of memory contention
Ingår i Parallel Architectures and Compilation Techniques - Conference Proceedings, PACT, s. 457-458, 2012
-
Cache Pirating: Measuring the Curse of the Shared Cache
Ingår i Proc. 40th International Conference on Parallel Processing, s. 165-175, 2011
-
Fast modeling of shared caches in multicore systems
Ingår i Proc. 6th International Conference on High Performance and Embedded Architectures and Compilers, s. 147-157, 2011
-
A simple statistical cache sharing model for multicores
Ingår i Proc. 4th Swedish Workshop on Multi-Core Computing, s. 31-36, 2011
-
Efficient software-based online phase classification
Ingår i International Symposium on Workload Characterization (IISWC'11), s. 104-115, 2011
-
StatStack: Efficient modeling of LRU caches
Ingår i Proc. International Symposium on Performance Analysis of Systems and Software, s. 55-65, 2010
-
Reducing Cache Pollution Through Detection and Elimination of Non-Temporal Memory Accesses
Ingår i Proc. International Conference for High Performance Computing, Networking, Storage and Analysis, s. 11, 2010
- DOI för Reducing Cache Pollution Through Detection and Elimination of Non-Temporal Memory Accesses
- Ladda ner fulltext (pdf) av Reducing Cache Pollution Through Detection and Elimination of Non-Temporal Memory Accesses
-
StatCC: a statistical cache contention model
Ingår i Proc. 19th International Conference on Parallel Architectures and Compilation Techniques, s. 551-552, 2010
-
A Software Technique for Reducing Cache Pollution
Ingår i Proc. 3rd Swedish Workshop on Multi-Core Computing, s. 59-62, 2010
-
Improving cache utilization using Acumem VPE
Ingår i Tools for High Performance Computing, s. 115-135, 2008
-
Conserving Memory Bandwidth in Chip Multiprocessors with Runahead Execution.
Ingår i 21st International Parallel and Distributed Processing Symposium, 2007
-
A case for low-complexity MP architectures
Ingår i Proc. Conference on Supercomputing, s. 559-570, 2007
-
A Statistical Multiprocessor Cache Model
Ingår i Proc. International Symposium on Performance Analysis of Systems and Software, s. 89-99, 2006
-
Modeling cache sharing on chip multiprocessor architectures
Ingår i Proc. International Symposium on Workload Characterization, s. 160-171, 2006
-
Multigrid and Gauss-Seidel smoothers revisited: Parallelization on chip multiprocessors
Ingår i Proc. 20th ACM International Conference on Supercomputing, s. 145-155, 2006
-
Exploiting Locality: A Flexible DSM Approach
Ingår i Proc. 20th IEEE International Parallel and Distributed Processing Symposium, 2006
-
TMA: A Trap-based Memory Architecture
Ingår i Proc. 20th ACM International Conference on Supercomputing, s. 259-268, 2006
-
Vasa: A Simulator Infrastructure with Adjustable Fidelity
Ingår i In Proceedings of the 17th IASTED International Conference on Parallel and Distributed Computing and Systems (PDCS 2005), Phoenix, Arizona, USA, November 2005., 2005
-
Exploring Processor Design Options for Java Based Middleware
Ingår i Proceedings of the 2005 International Conference on Parallel Processing (ICPP-05), 2005
-
Skewed Caches from a Low-Power Perspective
Ingår i Proceedings of Computing Frontiers, Ischia, Italy, May 2005, 2005
-
Exploiting Spatial Store Locality through Permission Caching in Software DSMs
Ingår i Proceedings of the 10th International Euro-Par Conference, s. 551, 2004
-
Bundling: Reducing the Overhead of Multiprocessor Prefetchers
Ingår i 18th International Parallel and Distributed Processing Symposium, 2004
-
StatCache: A Probabilistic Approach to Efficient and Accurate Data Locality Analysis
Ingår i 2004 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS-2004),, 2004
-
Hierarchical Backoff Locks for Nonuniform Communication Architectures
Ingår i Proceedings of the Ninth International Symposium on High Performance Computer Architecture (HPCA-9), Anaheim, California, USA, February 2003., 2003
-
THROOM — Supporting POSIX Multithreaded Binaries on a Cluster
Ingår i Euro-Par 2003, s. 760-769, 2003
-
Miss Penalty Reduction Using Bundled Capacity Prefetching in Multiprocessors
Ingår i Proceedings of the 17th InternationalParallel and Distributed Processing Symposium (IPDPS 2003), Nice, France, 2003
-
Memory System Behavior of Java-Based Middleware
Ingår i Proceedings of the Ninth International Symposium on High Performance Computer Architecture, 2003
-
RH Lock: A Scalable Hierarchical Spin Lock
Ingår i Proceedings of the 2nd Annual Workshop on Memory Performance Issues (WMPI 2002), held in conjunction with the 29th International Symposium on Computer Architecture (ISCA29), Anchorage, Alaska, USA, 2002
-
Efficient Synchronization for Non-Uniform Communication Architectures
Ingår i Proceedings of Supercomputing 2002, Baltimore, Maryland, USA, 2002
-
WildFire: A Scalable Path for SMPs
Ingår i Proc. Fifth Int. Symp. on High-Performance Computer Architecture, s. 172-181, 1999
Patent
-
Multiprocessing computer system employing capacity prefetching
2007
-
Computer system employing bundled prefetching
2007
-
System and method for reducing shared memory write overhead in multiprocessor systems
2006
-
Computer system including a promise array
2006
-
Multiprocessing computer system employing capacity prefetching
2006
-
Multiprocessing systems employing hierarchical back-off locks
2006
-
2005
-
2005
-
Multi-node computer system employing multiple memory response states
2005
-
2004
-
Multi-node computer system employing a reporting mechanism for multi-node transactions
2004
-
Multi-node computer system implementing global access state dependent transactions
2004
-
Multiprocessing computer system employing capacity prefetching
2004
-
Multi-node system with split ownership and access right coherence mechanism
2004
-
Performing virtual to global address translation in processing subsystem
2004
-
Multi-node system with global access states
2004
-
Multiprocessing systems employing hierarchical back-off locks
2004
-
2004
-
System and method for reducing shared memory write overhead in multiprocessor systems
2004
-
Computer system employing bundled prefetching
2004
-
Computer system including a promise array
2004
-
Multi-node computer system with proxy transaction to read data from a non-owning memory device
2004
-
Selective address translation in coherent memory replication
2003
-
2003
-
2003
-
2003
-
Multiprocessing systems employing hierarchical spin locks
2003
-
Communication error reporting mechanism in a multiprocessing computer system
2003
-
2002
-
Hybrid memory access protocol in a distributed shared memory computer system
2002
-
Hierarchical SMP computer System
2002
-
Skewed finite hashing function
2002
-
Selective address translation in coherent memory replication
2002
-
Shared memory system for symmetric microprocessor systems
2001
-
Multiprocessing system configured to perform efficient block copy operations
2001
-
2001
-
Multiprocessing system configured to perform efficient block copy operations
2001
-
Skewed finite hashing function
2001
-
Shared memory system for symmetric multiprocessor systems
2001
-
Communication error reporting mechanism in a multiprocessing computer system
2001
-
Selective address translation in coherent memory replication
2001
-
Communication error reporting mechanism in a multiprocessing computer system
2001
-
Multiprocessing system configured to perform efficient block copy operations
2001
-
Skewed finite hashing function
2001
-
Cache-less address translation
2001
-
Hybrid memory access protocol in a distributed shared memory computer system
2001
-
Method for increasing the speed of data processing in a computer system
2000
-
2000
-
2000
Rapporter
-
Delorean: Virtualized Directed Profiling for Cache Modeling in Sampled Simulation
2018
-
2017
-
The best of both works: A hybrid data-race-free cache coherence scheme
2017
-
A unified DVFS-cache resizing framework
2016
-
Perf-Insight: A Simple, Scalable Approach to Optimal Data Prefetching in Multicores
2015
-
Full Speed Ahead: Detailed Architectural Simulation at Near-Native Speed
2014
- Ladda ner fulltext (pdf) av Full Speed Ahead: Detailed Architectural Simulation at Near-Native Speed
-
Quantitative Characterization of Memory Contention
2012
-
Cache Pirating: Measuring the curse of the shared cache
2011
-
Multigrid and Gauss-Seidel smoothers revisited: Parallelization on chip multiprocessors
2006
-
TMA: A Trap-Based Memory Architecture
2005
-
Flexibility Implies Performance
2005
-
Adaptive Coherence Batching for Trap-Based Memory Architectures
2005
-
Low Power and Conflict Tolerant Cache Design
2004
-
Evaluation, Implementation and Performance of Write Permission Caching in the DSZOOM System
2004
-
StatCache: A Probabilistic Approach to Efficient and Accurate Data Locality Analysis
2003
-
The Elbow Cache: A Power-Efficient Alternative to Highly Associative Caches
2003
-
Low-Overhead Spatial and Temporal Data Locality Analysis
2003
-
THROOM: Running POSIX Multithreaded Binaries on a Cluster
2003
-
Bundling: Reducing the Overhead of Multiprocessor Prefetchers
2003
-
Latency-hiding and Optimizations of the DSZOOM Instrumentation System
2003