In VR the visor covers the upper part of the face. If the user is represented by an avatar that mimics the users facial expression, a camera cannot detect the upper part of the face. Train a NN that takes in the mouth/lower face parts and identifies the expression of the upper face part. Use videos of talking faces.
If you are interested in this topic or for more information please contact Prof. Helmut Hlavacs (helmut.hlavacs(at)univie.ac.at).