OpenSeeFace: Real-time Facial Landmark Tracking
GitHub Repo
BSD-2-Clause
July 20, 2026 at 08:22 PM
0 views

OpenSeeFace: Real-time Facial Landmark Tracking

@emilianavtProject Author

OpenSeeFace: Real-time Facial Landmark Tracking for Avatar Animation

[OSF.png embedded above]

Overview: What OpenSeeFace Is and How It Fits Into Your Avatar Pipeline

OpenSeeFace is a tracking library designed to detect facial landmarks in real time, with a clear purpose: it is not a stand-alone avatar puppeteering program. In practice, that means you can pair its tracking output with other tools to drive 3D avatars, 2D characters, or any other articulation system you’re building. A number of projects already leverage OpenSeeFace for avatar animation: VSeeFace uses OpenSeeFace to drive VRM models, while VTube Studio uses webcam-based tracking to animate Live2D assets. Even game engines like Godot can render OpenSeeFace-tracked data via dedicated plugins.

The core of OpenSeeFace is a facial landmark detection model built around MobileNetV3. To enable practical CPU-based inference on standard desktops, the model has been converted to ONNX, and the ONNX Runtime is used to achieve a tracking speed of roughly 30–60 frames per second for a single face. Four distinct models are provided, offering a spectrum of speed versus tracking quality so you can tailor performance to your hardware and application needs. The project’s name is a lighthearted pun on the “open seas” and the idea of “seeing faces,” with no deeper hidden meaning beyond the playful branding.

If you’re curious to see the default tracking model in action under various lighting and noise conditions, an up-to-date sample video is available, demonstrating the tracker’s robustness under real-world scenarios.

[OSF.png is included here to illustrate the project’s identity and branding]

Tracking quality: How OpenSeaFace Measures Up in Real World Use

One key design decision behind OpenSeeFace is that its landmark set is deliberately crafted to be useful for avatar animation rather than to be a gold-standard pixel-perfect face model. The landmarks resemble a 2D/3D face layout close to iBUG 68 points, but with important differences: two fewer points are placed at the mouth corners, and the contouring around the face uses quasi-3D contours rather than a strict visible outline. This choice yields a practical balance between tracking stability and the information needed to drive expressive avatars.

Because the focus is on reliable landmark placement for animation, numerical comparisons to other facial landmark systems (as found in academic literature) are not always apples-to-apples. What you can expect in practice is a tracker that maintains land­mark stability across a wide range of head poses, even when lighting is imperfect or the image resolution is not ideal. In particular, OpenSeeFace tends to hold onto landmark positions more stably than some alternatives when conditions are challenging, and it does a solid job representing a broad array of mouth shapes and expressions.

A few practical notes from experience and testing:

  • Eye landmarks are primarily useful to determine open/closed states and blink timing; their exact coordinates may be somewhat approximate, but that approximation is acceptable for blinking and gaze-driven avatar motion.
  • OpenSeeFace generally handles adverse conditions well—low light, noisy imagery, and reduced resolution—while keeping landmark positions stable over large head rotations.
  • In a direct comparison with MediaPipe, OpenSeeFace’s landmarks tend to stay steadier under difficult conditions and cover a wider variety of mouth poses. However, the eye region can be a bit less precise in some scenarios.

To illustrate the kinds of landmark outputs and the look of the detections, refer to the sample results images:

  • Landmarks visualization and detection results are shown in the Results series below.
  • A representative face-detection view is also provided to show how bounding boxes feed into the landmark estimation.

[EmiFace.png embedded here to show a robust face detection example and how a static bounding box can still yield strong landmark placement]

Usage: Getting OpenSeeFace Up and Running with Unity, Python, and Binaries

A practical Unity workflow is provided via a sample project that demonstrates VRM-based avatar animation driven by OpenSeeFace data. The face tracking itself is handled by a Python 3.7 script named facetracker.py. Because this is a command-line tool, you typically start it from a terminal (cmd on Windows) or you can automate its launch with a batch file.

If you’re using the released binaries on Windows, you can run the facetracker.exe inside the Binary folder, even if you don’t have Python installed. This convenience package assumes the corresponding models folder is located in the same directory as facetracker.exe (or in a shared parent folder). The binaries package also includes a version of ONNX Runtime without telemetry to keep privacy and performance in mind.

The Python-based tracker streams data via UDP, so you can run the tracker on one machine and consume its outputs on another, perhaps to offload computation while preserving the original camera feed for privacy reasons. Unity users can then attach the provided OpenSee component to receive UDP packets and expose the tracking data in a public field named trackingData. A companion OpenSeeShowPoints component visualizes the detected landmark points and serves as a handy example of how to render the data inside Unity.

For example, a basic demonstration scene can be created in Unity by placing an empty GameObject and attaching both the OpenSee and OpenSeeShowPoints components. While the scene is running, you can invoke the face tracker on a video file with a single command:

  • python facetracker.py --visualize 3 --pnp-points 1 --max-threads 4 -c video.mp4

If you installed dependencies via Poetry, you’ll want to run commands inside a Poetry shell (poetry shell) or prefix them with poetry run. This ensures the tracker publishes its own visualization while it streams tracking data to Unity.

The project also ships with an OpenSeeLauncher component that can start the face tracker program from within Unity. This launcher is designed to work with a PyInstaller-built executable distributed in the binary bundles. It exposes a small API to:

  • ListCameras(): returns the available camera names; the index corresponds to the cameraIndex field; -1 disables webcam capture.
  • StartTracker(): starts the tracker; if already running, it restarts with the current settings.
  • StopTracker(): stops the tracker; the tracker is also terminated automatically when the application closes or the launcher is destroyed. The launcher uses WinAPI job objects to ensure the tracker is properly terminated if the main application crashes, preventing orphan processes.

The OpenSeeIKTarget component can be used with IK systems such as FinalIK to animate head motion based on the detected landmarks, enabling more natural integration with various character rigs.

Expression detection: calibrating for expressive avatars

OpenSeeFace includes an OpenSeeExpression component that can be paired with OpenSeeFace to detect specific facial expressions. The system must be calibrated on a per-user basis, using a guided workflow inside the Unity Editor (or via equivalent public methods in the code). The calibration process typically involves capturing a set of example expressions, associating each with a name, and then training a detector to recognize those expressions in live data.

Calibration steps include:

  • Enter a name for the expression you want to calibrate.
  • Produce the expression and hold it while enabling recording.
  • Move your head in various directions while maintaining the expression, optionally pairing with talking for expressions meant to accompany speech.
  • Stop recording and repeat for additional expressions as needed.
  • Train and test the model, checking both detection accuracy and the live statistics displayed in the lower part of the component.
  • If any expression is underperforming, collect more data and re-train.
  • To remove data for an expression, simply enter its name and use the Clear option.
  • To save or load calibrated models and their training data, specify a filename with its full path and use Save/Load as appropriate.

Hints for effective calibration:

  • A practical target is about six expressions total, including neutral.
  • Warm up the tracker by making a few expressive faces and eyebrow movements before starting a full calibration session.
  • After obtaining a decent detection model, test all expressions in multiple poses and refine with additional data if needed.

General notes on performance and configuration

OpenSeeFace emphasizes robustness and practicality:

  • It remains robust under partial occlusion (glasses, masks, etc.) and imperfect lighting.
  • The model selection mechanism allows users to balance quality and speed:
  • The highest quality model is selected with --model 3.
  • The fastest model with the lowest tracking quality is --model 0.
  • Higher tracking quality tends to be less rigid and can complicate precise blinking and eyebrow motion capture.
  • CPU usage is a practical consideration: at 30 fps for a single face on a typical CPU, you’ll usually use less than one full core. If CPU load is excessive, dial back the frame rate (20 fps is often a safe target; values above 30 rarely justify the expense).
  • When tracking multiple faces, the system will run the face detector every --scan-every frames if you request more faces than are actually visible. To keep performance reasonable, cap the number of tracked faces to the actual count.

Models: four pretrained landmark configurations

OpenSeeFace ships with four pretrained landmark models, which can be selected with the --model flag. The fps values provided reflect running the model on a single face on a single CPU core; you can expect lower FPS if you increase the number of faces or if you run on hardware with fewer cores. The model lineup includes:

  • Model -1: An extremely fast, very low-accuracy model (213 fps, no gaze tracking).
  • Model 0: Very fast, low-accuracy model (68 fps).
  • Model 1: Slightly slower, better accuracy (59 fps).
  • Model 2: Slower, good accuracy (50 fps).
  • Model 3 (default): Slowest but highest accuracy (44 fps).

If you want to experiment with the more optimized options, you can locate additional weights in the project’s repository:

  • PyTorch-based weights for model.py are available via a Mega link.
  • Unoptimized ONNX models can be found in the Issues area of the repository.

Results: Visual Proof of Landmark Tracking

The project includes several visual samples illustrating landmark tracking outputs:

  • Results1.png and Results2.png show representative landmark placements and motion across frames.
  • Additional samples are provided as Results3.png and Results4.png to demonstrate how the tracker behaves under varying conditions and with different face orientations.
  • EmiFace.png offers a look at face detection stability and bounding-box performance across different scales and angles.

[Results1.png, Results2.png, Results3.png, Results4.png embedded here in sequence with captions]

Face detection: robust bounding boxes for reliable landmarks

The landmark model is designed to be robust to face size and orientation, allowing a relatively rough initial face bounding box to suffice for accurate landmark estimation. This yields a favorable speed-to-accuracy ratio that suits real-time avatar animation, where the priority is smooth, expressive motion rather than perfect face fitting in every frame.

Release builds: ready-to-run binaries with all dependencies

For users who want a quick start, the release builds bundle a facetracker.exe inside a Binary folder built with PyInstaller. To run smoothly, the models folder must be placed in the same directory as the executable (or in a common parent directory). Distributing the release should also include the Licenses folder to satisfy third-party library licensing terms. The release builds come with a custom ONNX Runtime build that excludes telemetry, in line with privacy-friendly design.

Dependencies and setup

OpenSeeFace’s Python-based workflow relies on a small set of well-known libraries. The official release notes list:

  • ONNX Runtime
  • OpenCV
  • Pillow
  • Numpy

You can install the core dependencies via pip:

  • pip install onnxruntime opencv-python pillow numpy

If you prefer a self-contained dependency environment, Poetry can be used to install all required packages into a dedicated virtual environment:

  • poetry install

This approach helps ensure reproducible results and avoids conflicts with other Python projects.

References and algorithmic underpinnings

OpenSeeFace draws from a mix of modern neural network design ideas and established face-alignment research:

  • The landmark detection model is influenced by MobileNetV3 and its search for efficient neural architectures suitable for real-time inference on CPU hardware.
  • The training leverages Adaptive Wing Loss for robust landmark heatmap regression, a technique designed to improve keypoint localization under challenging conditions.
  • The face detection component uses a heatmap regression approach, with RetinaFace-based architectures serving as a reference implementation for robust single-stage localization.
  • A number of public datasets contribute to the training and evaluation pipeline, including LS3D-W (for 66-point landmarks), WFLW (for extended fitting, with adaptation to 66 points), MPIIGaze (for gaze and blink detection), and UnityEyes (for synthetic eye rendering).
  • The algorithmic lineage also includes widely cited architectures like U-Net and MobileNet, and optimization strategies for real-time face tracking in interactive environments.

The project’s references section includes explicit citations to the literature underpinning the model design, as well as links to the datasets and tools used during development:

  • Training datasets: LS3D-W, WFLW, MPIIGaze, UnityEyes, and synthetic data augmentations.
  • Core algorithms and architectures: MobileNets, U-Net, RetinaFace, and the broader context of real-time landmark tracking.
  • Expression detection: LIBSVM is used in some expression components for classification tasks, while eye gaze and blink signals are derived from landmark positions.

A note on licensing and distribution

OpenSeeFace is distributed under the BSD 2-Clause license. The release binaries carry the necessary licenses for third-party libraries included in the package. If you redistribute, ensure you provide the Licenses folder to comply with licensing requirements.

Ackowledgments and community

The author thanks a number of testers and collaborators who helped refine the project, including social media contributors and fellow developers. The project’s collaborative nature is reflected in the community-driven testing and feedback that shaped both the tracking quality and the usability features described in this post.

What’s next: practical tips for using OpenSeeFace in production

  • Start with the default model (Model 3) to gauge baseline performance on your hardware, then experiment with faster models if you need more headroom for desktop or laptop workloads.
  • When streaming data to Unity, consider running the tracker on a dedicated PC or virtual machine to minimize latency and protect privacy by separating the camera feed from the avatar’s rendering pipeline.
  • If you work with multiple faces, adjust the --faces setting so the detection model isn’t unnecessarily invoked on frames where more faces aren’t present; this helps preserve CPU headroom.
  • Use the expression calibration workflow to tailor expression detection to your character and performance style, ensuring natural lip-sync, eyebrow motion, and other facial cues line up with your animation rig.
  • Keep an eye on frame rate targets. Real-time avatar animation benefits from stable frame rates; avoid chasing ultra-high fps if your CPU or GPU can’t sustain it.
  • OSF.png: The project branding and identity
  • EmiFace.png: A visualization of face detection and bounding boxes
  • Results1.png, Results2.png, Results3.png, Results4.png: Landmark tracking visualizations across multiple samples
  • Additional visuals: The embedded images illustrate how landmark points map to facial regions and how those points feed into Unity and other rendering pipelines

[EmiFace.png and Results1-4.png embedded here, with descriptive captions]

Conclusion: OpenSeeFace as a Practical Bridge Between Vision and Expression

OpenSeeFace embodies a pragmatic approach to real-time facial landmark tracking: it emphasizes robust, avatar-friendly outputs, flexible deployment options (binaries, Python-based pipelines, and Unity integrations), and a scalable model lineup that lets you tune speed and quality to fit your project. It’s not a complete avatar engine, but it provides a reliable, fast, and adaptable foundation for driving expressive 3D or 2D characters from live facial data.

If you’re building an interactive experience, a streaming application, or a cinematic toolset that requires believable facial animation driven by webcam or video input, OpenSeeFace offers a well-documented, community-supported path forward. The combination of lightweight MobileNetV3-based landmarks, ONNX Runtime acceleration, and a robust Unity integration makes it a compelling choice for developers who value both performance and flexibility.

[Images for reference]

  • OSF.png
  • EmiFace.png
  • Results1.png
  • Results2.png
  • Results3.png
  • Results4.png

Endnote: All the pieces you need are in the project repository—sample Unity scenes, Python scripts, release binaries, and a full set of documentation that covers installation, usage, and extension possibilities. Whether you’re prototyping a concept or shipping a production character with real-time facial animation, OpenSeeFace is designed to help you turn expressive faces into lively digital performances.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
openseeface
Created
July 20
Last Updated
July 16, 2026 at 11:54 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.