Gradient Regularization
- Gradient harmonization: [1]
[1] Gradient Harmonized Single-stage Detector, AAAI, 2019
[1] Gradient Harmonized Single-stage Detector, AAAI, 2019
Geometry feature generation based on unsupervisely detected landmarks. [1]
Disentangle bottleneck features into category-invariant features and category-specific features. Category-invariant features encode the pose information.

Corneal reflection-based methods
Appearance based methods
Calibration: obtain the visual axis and kappa angle for each person.
Facial landmarks detection
Head Pose Estimation
[MPIIGaze]: fine-grained annotation
[Eyediap]: RGB-D
Distinguish generated fake images and real images in the freqency domain. [2]
An image can be composed of or decomposed into low-frequency part and high-frequency part [3] [8] [4] [10]
Kai Xu, Minghai Qin, Fei Sun, Yuhao Wang, Yen-Kuang Chen, Fengbo Ren, “Learning in the Frequency Domain”, CVPR, 2020.
Wang, Sheng-Yu, et al. “CNN-generated images are surprisingly easy to spot… for now.” arXiv preprint arXiv:1912.11035 (2019).
ayush Bansal, Yaser Sheikh, Deva Ramanan, “PixelNN: Example-based Image Synthesis”, ICLR 2018.
Yanchao Yang, Stefano Soatto, “FDA: Fourier Domain Adaptation for Semantic Segmentation”, CVPR 2020.
Roy, Hiya, et al. “Image inpainting using frequency domain priors.” arXiv preprint arXiv:2012.01832 (2020).
Shen, Xing, et al. “DCT-Mask: Discrete Cosine Transform Mask Representation for Instance Segmentation.” arXiv preprint arXiv:2011.09876 (2020).
Suvorov, Roman, et al. “Resolution-robust Large Mask Inpainting with Fourier Convolutions.” WACV (2021).
Yu, Yingchen, et al. “WaveFill: A Wavelet-based Generation Network for Image Inpainting.” ICCV, 2021.
Mardani, Morteza, et al. “Neural ffts for universal texture image synthesis.” NeurIPS (2020).
Cai, Mu, et al. “Frequency domain image translation: More photo-realistic, better identity-preserving.” ICCV, 2021.
Predict visual feature of one future frame [1]
Predict optical flow of one future frame [2]
Predict one future frame [4] (a special case of video prediction)
Predict future trajectories [5]
Predict optical flows of future frames, and then obtain future frames [3]
Vondrick, Carl, Hamed Pirsiavash, and Antonio Torralba. “Anticipating visual representations from unlabeled video.” CVPR, 2016.
Gao, Ruohan, Bo Xiong, and Kristen Grauman. “Im2flow: Motion hallucination from static images for action recognition.” CVPR, 2018.
Li, Yijun, et al. “Flow-grounded spatial-temporal video prediction from still images.” ECCV, 2018.
Xue, Tianfan, et al. “Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks.” NIPS, 2016.
Walker, Jacob, et al. “An uncertain future: Forecasting from static images using variational autoencoders.” ECCV, 2016.
Meta-learning method: [1]
Delta-based: delta between each pair of samples [2]; delta between each sample and class center [3] [4]
[1] Zhang, Ruixiang, et al. “Metagan: An adversarial approach to few-shot learning.” NIPS, 2018.
[2] Schwartz, Eli, et al. “Delta-encoder: an effective sample synthesis method for few-shot object recognition.” Advances in Neural Information Processing Systems. 2018.
[3] Liu, Jialun, et al. “Deep Representation Learning on Long-tailed Data: A Learnable Embedding Augmentation Perspective.” arXiv preprint arXiv:2002.10826 (2020).
[4] Yin, Xi, et al. “Feature transfer learning for face recognition with under-represented data.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2019.
Light-weighted network structure
SqueezeNet, MobileNet, and ShuffleNet share the same idea: decouple the temporal convolution and spatial convolution to reduce the nummber of parameters, sharing the similar spirit with Pseudo-3D Residual Networks. SqueezeNet is serial while MobileNet and ShuffleNet are parrallel. MobileNet is a special case of ShuffleNet when using only one group.
Low-rank approximation ($k\times k \times c\times d = k\times k\times c\times d’ + 1\times 1\times d’\times d$) also falls into the above scope. The difference between MobileNet and Low-rank approximation is layerwise convolution or not.
Tweak network structure
Compress weights
Computation
Sparsity regularization
Efficient Inference
Good introduction slides: http://cs231n.stanford.edu/slides/2017/cs231n_2017_lecture15.pdf