ovrgpuprofiler is a performance monitoring CLI tool for Meta Quest headsets that developers can use to access a range of real-time GPU metrics and perform render stage traces. ovrgpuprofiler is included with the Meta Quest runtime and does not need to be manually installed.ovrgpuprofiler. If not using a shell, precede all commands in this topic with adb shell <command>.ovrgpuprofiler -m
47 metrics supported:
1 Clocks / Second
2 GPU % Bus Busy
3 % Vertex Fetch Stall
4 % Texture Fetch Stall
5 L1 Texture Cache Miss Per Pixel
6 % Texture L1 Miss
7 % Texture L2 Miss
8 % Stalled on System Memory
9 Pre-clipped Polygons/Second
10 % Prims Trivially Rejected
11 % Prims Clipped
ovrgpuprofiler prints, starting at 1. They are not fixed hardware counter identifiers, so the ID for a given metric can differ across devices and runtime versions. List the metrics on the device you are profiling and read the IDs from that output. ovrgpuprofiler rejects an ID outside the printed range with a message naming the valid range.ovrgpuprofiler -m -v provides the same list with more verbose descriptions for each metric.ovrgpuprofiler -r<metric ID number>
ovrgpuprofiler -r4 and the console prints data every second until you press Ctrl-C.ovrgpuprofiler -r"4,5,6". The following shows output from ovrgpuprofiler -r"4,5,6":$ ovrgpuprofiler -r"4,5,6" % Texture Fetch Stall : 2.449 L1 Texture Cache Miss Per Pixel : 0.124 % Texture L1 Miss : 20.338 % Texture Fetch Stall : 2.369 L1 Texture Cache Miss Per Pixel : 0.122 % Texture L1 Miss : 20.130 % Texture Fetch Stall : 2.580 L1 Texture Cache Miss Per Pixel : 0.127 % Texture L1 Miss ...
ovrgpuprofiler supports render stage GPU tracing on a tile-per-tile level. Unlike direct-mode GPUs, which execute draw calls sequentially, tile-based renderers batch draw calls for an entire surface. That surface is then split into tiles that are computed sequentially, where each tile executes all the draw calls that touched that tile. ovrgpuprofiler can tell you how much time was spent in each rendering stage for each surface rendered during a trace’s duration.ovrgpuprofiler -e
ovrgpuprofiler -e com.YourCompany.YourApp. Without an argument, the mode applies to every app started afterward.ovrgpuprofiler -i shows whether detailed GPU profiling mode is currently enabled. Use ovrgpuprofiler -d to disable it.ovrgpuprofiler -t
-t argument. For example, ovrgpuprofiler -t1.2 would run a trace for 1.2 seconds. A duration that cannot be parsed falls back to the 100 ms default and prints a warning.-c to keep polling the trace until you press Ctrl-C, which reports results in batches instead of holding a single long capture in memory. Add -l to run the trace in low-overhead mode, which prints one line per surface and omits the per-stage and per-bin breakdown in exchange for a more accurate measurement.Process com.YourCompany.YourApp has 24 surface executions Surface 1 | 1216x1344 | color 32bit, depth 24bit, stencil 0 bit, MSAA 4, Mode: 1 (HwBinning) | 60 128x224 bins ( 60 rendered) | 5.08 ms | 130 stages : Binning : 0.623ms Render : 1.877ms StoreColor : 0.309ms Blit : 0.002ms Preempt : 1.286ms
ovrgpuprofiler prints the mode number followed by a short name:unknown render mode <number>.Compute 4 | | 1.42 ms
-v to expand each surface into the bins and stages that make it up:ovrgpuprofiler -t1 -v
Surface 1 | 1216x1344 | color 32bit, depth 24bit, stencil 0 bit, MSAA 4, Mode: 1 (HwBinning) | 60 128x224 bins ( 60 rendered) | 5.08 ms | 130 stages : Binning : 0.623ms Render : 1.877ms StoreColor : 0.309ms
Bin 0 | topLeft 0 x0 | 1x1 logical bins | Fov 8/8
Stage 0 : Binning 0.623 ms
Stage 1 : Render 0.031 ms
Stage 2 : StoreColor 0.005 ms
Bin 1 | topLeft 128 x0 | 1x1 logical bins | Fov 8/8
Stage 3 : Render 0.029 ms
Stage 4 : StoreColor 0.005 ms
Fov L<x>/<y> R<x>/<y>.ovrgpuprofiler outputs one surface line per slice for multiview apps. This means that there is one surface for each eye. You must add the render times of two eye surfaces for the total frame time.ovrgpuprofiler outputs one surface line for both views of the surface, due to how the Adreno 650v3 GPU processes multiview commands (Hardware Multiview). On Meta Quest 2, bins of multiview surfaces are shared between both views. Therefore, the following output:135 96x176 bins
135 96x176x2 bins
ovrgpuprofiler outputs one surface line per eye (as on Meta Quest) or one surface line for both views (as on Meta Quest 2 with Hardware Multiview) has not been verified for these devices. Refer to the release notes for your runtime version or contact developer support for confirmation.ovrgpuprofiler can sample GPU metrics for each render stage in a trace. This attributes counter values to a specific binning, render, or store stage of a specific surface, rather than to the frame as a whole.-m with -t to list the trace metrics supported on the connected device:ovrgpuprofiler -m -t
-v for verbose descriptions. As with the other catalogs, the IDs are positions in this list and vary by device and runtime version.-s (--renderstage-metrics) and combine it with -t:ovrgpuprofiler -t3 -s "1,2,3"
-s requires -t. Running -s on its own reports an error and exits with a non-zero status, as does an empty or fully invalid ID list. Detailed GPU profiling mode must be enabled first, the same as for any render stage trace.Tracing for 3.0 seconds...
Captured 2 metrics, 1 rejected by SoC counter budget
rejected: % Texture L2 Miss
Process com.YourCompany.YourApp has 24 surface executions
Surface 1 | 1216x1344 | color 32bit, depth 24bit, stencil 0 bit, MSAA 4, Mode: 1 (HwBinning) | 60 128x224 bins ( 60 rendered) | 5.08 ms | 130 stages : Binning : 0.623ms Render : 1.877ms StoreColor : 0.309ms
Bin 0 | topLeft 0 x0 | 1x1 logical bins | Fov 8/8
Stage 0 : Binning 0.623 ms
Clocks : 58912.000
% Shaders Busy : 39.907
Stage 1 : Render 0.031 ms
Clocks : 3204.000
% Shaders Busy : 61.442
-s expands the per-stage breakdown on its own, so -v is not required to see the stage lines. A metric the counter allocator rejected is absent from the per-stage values.rejected list. Run several traces with smaller subsets when you need more metrics than one pass holds.-s with -x is not supported: per-draw mode takes precedence, -s is ignored, and the tool prints a warning.ovrgpuprofiler can sample GPU metrics on a per-draw-call basis within each surface. This is useful for identifying which draw calls dominate a frame’s GPU cost or for comparing the relative cost of individual draws.ovrgpuprofiler -e and restarting the target app.ovrgpuprofiler -x -m
-v for more verbose descriptions. As with real-time metrics, the available set varies by device and runtime version.-x (draw call) flag with -t (trace), passing a comma-separated list of per-draw metric IDs to record:ovrgpuprofiler -t3 -x="1,2,3,4,12,15,16,17,18,22,27,28,29,38,49,50"
-t3 runs a 3-second trace (the default trace length is 0.1 seconds) and the -x argument selects 16 per-draw metric IDs to sample. A per-draw trace with no metric IDs reports that there is nothing to track and stops.Tracing for 3.0 seconds...
Process com.YourApp has 30 commandbuffers
Frame 1 : Draw 1 Label 0xffffffff
LRZ State: TestEnabled, WriteEnabled <0x03>
Clocks : 58912.000
% Vertex Fetch Stall : 0.061
% Shaders Busy : 39.907
Fragments Shaded : 3.203
Frame 1 : Draw 2 Label 0xffffffff
LRZ State: TestEnabled, WriteEnabled <0x03>
Clocks : 44008.000
% Vertex Fetch Stall : 0.000
...
TestEnabled, TestEnabled, WriteEnabled, or Disabled. A driver that does not report the state prints Unknown(old driver?).ovrgpuprofiler -d.ovrgpuprofiler:| Argument | Description |
|---|---|
-h/--help
| Prints the list of supported arguments and exits.
|
-r/--realtime
| Prints the value of the real-time metrics every second. Accepts an optional comma-separated list of metrics IDs to track.
|
-m/--metrics
| Prints the list of available real-time metrics IDs, their name, and their description. When combined with -x, lists the per-draw metric catalog instead; when combined with -t, lists the render-stage trace metric catalog.
|
-v/--verbose
| Adds more detailed information to most other commands. With -t, expands each surface into its per-bin and per-stage breakdown.
|
-e/--enable-detailed
| Enables detailed profiling mode on the GPU driver; required for render stage tracing. Only applies to applications started after this mode is started. Accepts an optional application name to attach to that application only.
|
-d/--disable-detailed
| Disables detailed profiling mode on the GPU driver.
|
-i/--is-detailed
| Queries if the GPU driver is in detailed profiling mode.
|
-t/--trace
| Executes a render stage trace, with an optional trace length as argument in seconds (default 0.1 s).
|
-c/--continuous
| If you specify this along with -t/--trace, the results of the render stage trace are polled periodically to reduce memory pressure.
|
-l/--low-overhead
| If you specify this along with -t/--trace, the render stage trace is performed in low-overhead mode, which omits many details in exchange for a more accurate measurement.
|
-s/--renderstage-metrics
| Captures per-render-stage metric values during a trace. Takes a required comma-separated list of render-stage metric IDs, for example -s "1,2,3" or --renderstage-metrics="1,2,3". Must be combined with -t; ignored when -x is also present.
|
-x/--drawcall
| Switches the tool into per-draw mode. Combine with -m to list per-draw metrics, or with -t and an "=" comma-separated list of metric IDs to record a per-draw trace (for example, -x="1,2,3").
|