The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more com...