Develop

Stereo with foveated inset (quad views)

Updated: Sep 14, 2026
XR_VIEW_CONFIGURATION_TYPE_PRIMARY_STEREO_WITH_FOVEATED_INSET is a view configuration that renders four views per frame instead of the usual two: a wide, low-resolution outer view per eye, plus a narrow, high-resolution inset view per eye. The inset views cover a small portion of the field of view at higher pixels per degree, while the outer views cover the full field of view at lower pixels per degree. On devices with eye tracking, the runtime steers the inset to follow the user’s gaze each frame.
Because the human visual system only resolves high detail in the foveal region, this configuration can render a significantly smaller total pixel count while preserving perceived sharpness where the user is actually looking. Splitting the per-eye image into a high-density inset and a lower-density outer view typically cuts the total pixels rendered by roughly 50% compared with a uniform-density stereo target at matching foveal resolution.
Per-eye coverage diagram (shown post-barrel-distortion): the red region is the wide-FOV outer view at lower pixels per degree; the green region is the narrow-FOV high-density inset overlaid on top. The outer view still renders the full field of view including the area the inset covers.
The runtime only advertises this view configuration when your application opts in via the AndroidManifest.xml declaration described below. Without the opt-in, xrEnumerateViewConfigurations will not return it, even on devices that support it.

How it works

When this view configuration is active, xrEnumerateViewConfigurationViews returns four entries in the following fixed order:
IndexView
0
Left outer (wide field of view, lower pixels per degree)
1
Right outer (wide field of view, lower pixels per degree)
2
Left inset (narrow field of view, higher pixels per degree)
3
Right inset (narrow field of view, higher pixels per degree)
The inset and outer views for the same eye share the same pose returned by xrLocateViews — they are co-located cameras with different fields of view. The runtime guarantees that the inset’s projected rectangle is fully contained within the outer view’s projected rectangle, so the outer view must still render the full field of view including the region the inset covers. The compositor blends between the inset and outer images at the inset’s edges to hide the resolution boundary.
When eye tracking is active, the inset’s fov returned from xrLocateViews shifts each frame to follow the user’s gaze (the runtime averages the gaze direction across both eyes so that the left and right insets remain aligned). When eye tracking is not active, the inset is fixed at the center of the field of view. See Eye tracking interaction below for the conditions under which eye tracking is active.

Enabling the view configuration

AndroidManifest.xml

Add the following <uses-feature> tag inside the <manifest> element of your AndroidManifest.xml:
<uses-feature
    android:name="com.oculus.feature.QUAD_VIEWS"
    android:required="false" />
Use android:required="false" if your application falls back to XR_VIEW_CONFIGURATION_TYPE_PRIMARY_STEREO when the inset configuration is not available. Use android:required="true" only if your application cannot run without it — note that this restricts installation to Quest devices on Horizon OS v85 or higher.
Without this manifest tag, the runtime omits STEREO_WITH_FOVEATED_INSET from the list returned by xrEnumerateViewConfigurations, regardless of device capability.

Implementation

Step 1: Enumerate and select the view configuration

When STEREO_WITH_FOVEATED_INSET is supported, the Meta runtime returns it first from xrEnumerateViewConfigurations, signaling that it is the runtime’s preferred configuration. Following the standard “select the first supported” pattern automatically opts your application in:
uint32_t viewConfigCount = 0;
xrEnumerateViewConfigurations(instance, systemId, 0, &viewConfigCount, NULL);

XrViewConfigurationType viewConfigs[8];
xrEnumerateViewConfigurations(instance, systemId, viewConfigCount,
    &viewConfigCount, viewConfigs);

XrViewConfigurationType selectedConfig = viewConfigs[0];
// selectedConfig will be XR_VIEW_CONFIGURATION_TYPE_PRIMARY_STEREO_WITH_FOVEATED_INSET
// if the manifest tag is set and the device supports it; otherwise PRIMARY_STEREO.

Step 2: Query the four view dimensions

Call xrEnumerateViewConfigurationViews with the selected configuration to retrieve the per-view recommended dimensions:
uint32_t viewCount = 0;
xrEnumerateViewConfigurationViews(instance, systemId, selectedConfig,
    0, &viewCount, NULL);
// viewCount == 4 for STEREO_WITH_FOVEATED_INSET

XrViewConfigurationView configViews[4] = {
    {XR_TYPE_VIEW_CONFIGURATION_VIEW},
    {XR_TYPE_VIEW_CONFIGURATION_VIEW},
    {XR_TYPE_VIEW_CONFIGURATION_VIEW},
    {XR_TYPE_VIEW_CONFIGURATION_VIEW},
};
xrEnumerateViewConfigurationViews(instance, systemId, selectedConfig,
    viewCount, &viewCount, configViews);

// configViews[0..1] = outer views, configViews[2..3] = inset views.
// The recommended dimensions differ between outer and inset views.

Step 3: Allocate swapchains for four views

The recommended layout is two multiview swapchains — one for the two outer views (arraySize = 2, sized to the outer recommended dimensions) and one for the two inset views (arraySize = 2, sized to the inset recommended dimensions). This pairs naturally with the recommended draw order in Step 5. The recommended dimensions for outer and inset views may differ (currently they are the same — the inset still achieves a higher pixels-per-degree because it covers a smaller field of view in the same pixel budget — but this can change in future runtime versions, so always size each pair from the values returned by xrEnumerateViewConfigurationViews).
// Outer views (indices 0 and 1) share dimensions and use one multiview swapchain.
XrSwapchainCreateInfo outerCreateInfo = {
    .type = XR_TYPE_SWAPCHAIN_CREATE_INFO,
    .usageFlags = XR_SWAPCHAIN_USAGE_COLOR_ATTACHMENT_BIT |
                  XR_SWAPCHAIN_USAGE_SAMPLED_BIT,
    .format = vulkanColorFormat,
    .sampleCount = 1,
    .width  = configViews[0].recommendedImageRectWidth,
    .height = configViews[0].recommendedImageRectHeight,
    .faceCount = 1,
    .arraySize = 2,  // multiview, two outer eyes
    .mipCount = 1,
};
XrSwapchain outerSwapchain;
xrCreateSwapchain(session, &outerCreateInfo, &outerSwapchain);

// Inset views (indices 2 and 3) share dimensions and use one multiview swapchain.
XrSwapchainCreateInfo insetCreateInfo = outerCreateInfo;
insetCreateInfo.width  = configViews[2].recommendedImageRectWidth;
insetCreateInfo.height = configViews[2].recommendedImageRectHeight;
XrSwapchain insetSwapchain;
xrCreateSwapchain(session, &insetCreateInfo, &insetSwapchain);
A single arraySize = 4 swapchain is only viable when all four recommended dimensions match. Even then, two multiview-2 swapchains are still preferred — see Step 5 for why batching the outer pair and inset pair as separate passes is the right rendering boundary.

Step 4: Locate four views each frame

Pass the four-view configuration to xrLocateViews with capacity 4. The returned views[2..3].fov will shift each frame on eye-tracked devices:
XrView views[4] = {
    {XR_TYPE_VIEW}, {XR_TYPE_VIEW}, {XR_TYPE_VIEW}, {XR_TYPE_VIEW},
};
XrViewState viewState = {XR_TYPE_VIEW_STATE};
XrViewLocateInfo locateInfo = {
    .type = XR_TYPE_VIEW_LOCATE_INFO,
    .viewConfigurationType =
        XR_VIEW_CONFIGURATION_TYPE_PRIMARY_STEREO_WITH_FOVEATED_INSET,
    .displayTime = predictedDisplayTime,
    .space = appSpace,
};

uint32_t viewCount = 4;
xrLocateViews(session, &locateInfo, &viewState, 4, &viewCount, views);

Step 5: Render the views as two multiview passes

Render the two outer views together as a single multiview-2 pass into the outer swapchain, then render the two inset views together as a separate multiview-2 pass into the inset swapchain. Do not use multiview-4 — pairing the outer eyes together and the inset eyes together is the right batching boundary because:
  • The outer pair covers the full field of view; the inset pair covers a smaller field of view. Their visible object sets, LOD selection, and (where used) per-view render areas differ. Treating them as two separate multiview-2 passes lets you tune each pair independently — for example, drawing fewer objects in the inset, or biasing inset LODs toward higher detail — without paying multiview-4’s cost across views whose visible content overlaps very little.
  • When the outer and inset recommended dimensions differ (which they may, depending on runtime version), a single multiview-4 swapchain is not even possible — the two pairs require separate swapchains.
Two important constraints from the OpenXR specification apply across both passes:
  • The outer views must cover the full field of view, including the region covered by the inset. Do not skip rendering the inner region of the outer views — the compositor uses the outer view to fill outside the inset and to blend at the inset’s edge.
  • The inset views’ pose always equals the corresponding outer eye’s pose. Only the fov differs (and changes per frame when eye tracking is active).

Step 6: Submit a four-view projection layer

Submit a single XrCompositionLayerProjection containing all four projection views with viewCount = 4. Views 0 and 1 reference the outer multiview swapchain (array indices 0 and 1); views 2 and 3 reference the inset multiview swapchain (array indices 0 and 1):
XrCompositionLayerProjectionView projViews[4];

// Outer views: array indices 0 and 1 of outerSwapchain.
for (int i = 0; i < 2; i++) {
    projViews[i] = (XrCompositionLayerProjectionView){
        .type = XR_TYPE_COMPOSITION_LAYER_PROJECTION_VIEW,
        .pose = views[i].pose,
        .fov  = views[i].fov,
        .subImage = {
            .swapchain = outerSwapchain,
            .imageArrayIndex = i,
            .imageRect = {
                .offset = {0, 0},
                .extent = {
                    configViews[i].recommendedImageRectWidth,
                    configViews[i].recommendedImageRectHeight,
                },
            },
        },
    };
}

// Inset views: array indices 0 and 1 of insetSwapchain.
for (int i = 0; i < 2; i++) {
    projViews[i + 2] = (XrCompositionLayerProjectionView){
        .type = XR_TYPE_COMPOSITION_LAYER_PROJECTION_VIEW,
        .pose = views[i + 2].pose,
        .fov  = views[i + 2].fov,
        .subImage = {
            .swapchain = insetSwapchain,
            .imageArrayIndex = i,
            .imageRect = {
                .offset = {0, 0},
                .extent = {
                    configViews[i + 2].recommendedImageRectWidth,
                    configViews[i + 2].recommendedImageRectHeight,
                },
            },
        },
    };
}

XrCompositionLayerProjection projLayer = {
    .type = XR_TYPE_COMPOSITION_LAYER_PROJECTION,
    .space = appSpace,
    .viewCount = 4,
    .views = projViews,
};

Eye tracking interaction

When eye tracking is active, the runtime samples the user’s gaze each frame and steers the inset views to follow it. The runtime averages the gaze direction across the two eyes so that the left and right insets remain aligned with each other, which keeps stereo correspondence intact at the inset boundary.
Eye tracking is only active when both of the following are true:
  • The user has granted the eye tracking permission to your application.
  • Your application has enabled an OpenXR extension that requires eye tracking, such as XR_EXT_eye_gaze_interaction⁠.
If either condition is not met, the inset is fixed at the center of the field of view and xrLocateViews simply returns a stable views[2..3].fov each frame. No additional adaptation is required from the application.

Compatibility

  • The application must support a fallback to XR_VIEW_CONFIGURATION_TYPE_PRIMARY_STEREO for devices and runtime configurations where the inset configuration is not advertised. Use android:required="false" on the manifest tag so that your application can install on every Quest device.
  • The runtime advertises the view configuration only when both the manifest opt-in (or debug setprop) is present and the device supports it.

Combining with other rendering optimizations

Do not combine with fixed foveated rendering (FFR)

Avoid using fixed foveated rendering on top of STEREO_WITH_FOVEATED_INSET. The inset/outer split is itself a foveation technique — applying a fragment density map to the outer view on top of an already-reduced-density buffer compounds aliasing in the periphery without meaningful additional savings, and the inset is small enough that FFR provides little benefit there either. Use one or the other, not both.

When combining with symmetric projection, use per-view render areas

If your application uses symmetric projection for the outer views, prefer the VK_QCOM_multiview_per_view_render_areas Vulkan extension to limit the rendered pixels to each eye’s actual visible region. With symmetric projection, the symmetric outer buffer is wider than each eye’s true field of view; per-view render areas let you skip shading the off-screen tiles directly. This is a better fit for quad views than relying on a fragment density map to discard those tiles, because it avoids the FFR/inset interaction described above.

Tips and best practices

  • Treat the outer views as a full coverage pass. Do not optimize by skipping the inner region — the compositor relies on the outer view to fill the area immediately outside the inset and to blend the boundary.
  • Allocate inset and outer swapchains at the recommended sizes from xrEnumerateViewConfigurationViews. The runtime sizes them so that the inset is at higher pixels per degree than the outer view; do not equalize them.
  • Cull the inset from the already-culled outer. The inset’s frustum is fully contained within the outer view’s frustum, so once you have the outer view’s culling result, you can derive the inset’s visible-object set by further culling that smaller list against the inset frustum — you do not need to run a full broad-phase cull twice per eye.
  • Use different LOD selection for the inset and outer views. The inset is the foveal region and benefits from higher-detail LODs; the outer view is peripheral and can use coarser LODs. Bias your LOD selection per view rather than rendering both with the same LODs derived from the eye-to-object distance.
  • Profile carefully. The pixel reduction translates into shader time savings, but adds fixed overhead per additional view (extra draw call replication, view-matrix work, additional culling pass). Fragment-heavy scenes — heavy shaders, complex post-processing, expensive material work — typically see the largest net wins; very lightweight scenes may see less benefit or none.