OPEN Active inference under visuoproprioceptive conflict: Simulation and empirical results
OPEN Active inference under visuoproprioceptive conflict: Simulation and empirical results
It has been suggested that the brain controls hand movements via internal models that rely on visual and proprioceptive cues about the state of the hand. In active inference formulations of such models, the relative influence of each modality on action and perception is determined by how precise (reliable) it is expected to be. The 'top-down' affordance of expected precision to a particular sensory modality is associated with attention. Here, we asked whether increasing attention to (i.e., the precision of) vision or proprioception would enhance performance in a hand-target phase matching task, in which visual and proprioceptive cues about hand posture were incongruent. We show that in a simple simulated agent-based on predictive coding formulations of active inference-increasing the expected precision of vision or proprioception improved task performance (target matching with the seen or felt hand, respectively) under visuoproprioceptive conflict. Moreover, we show that this formulation captured the behaviour and self-reported attentional allocation of human participants performing the same task in a virtual reality environment. Together, our results show that selective attention can balance the impact of (conflicting) visual and proprioceptive cues on action-rendering attention a key mechanism for a flexible body representation for action.
Controlling the body's actions in a constantly changing environment is one of the most important tasks of the human brain. The brain solves the complex computational problems inherent in this task by using internal probabilistic (Bayes-optimal) models. These models allow the brain to flexibly estimate the state of the body and the consequences of its movement, despite noise and conduction delays in the sensorimotor apparatus, via iterative updating by sensory prediction errors from multiple sources. The state of the hand, in particular, can be informed by vision and proprioception. Here, the brain makes use of an optimal integration of visual and proprioceptive signals, where the relative influence of each modality-on the final 'multisensory' estimate-is determined by its relative reliability or precision, depending on the current context.
These processes can be investigated under an experimentally induced conflict between visual and proprioceptive information. The underlying rationale here is that incongruent visuoproprioceptive cues about hand position or posture have to be integrated (provided the incongruence stays within reasonable limits), because the brain's body model entails a strong prior belief that information from both modalities is generated by one and the same external cause; namely, one's hand. Thus, a partial recalibration of one's unseen hand position towards the position of a (fake or mirror-displaced) hand seen in an incongruent position has been interpreted as suggesting an (attempted) resolution of visuoproprioceptive conflict to maintain a prior body representation.
Importantly, spatial or temporal perturbations can be introduced to visual movement feedback during action-by displacing the seen hand position in space or time, using video recordings or virtual reality. Such experiments suggest that people are surprisingly good at adapting their movements to these kinds of perturbations; i.e., they adjust to novel visuomotor mappings by means of visuoproprioceptive recalibration or adaptation. During motor tasks involving the resolution of a visuoproprioceptive conflict, one typically observes increased activity in visual and multisensory brain areas. The remapping required for this resolution is thought to be augmented by attenuation of proprioceptive cues. The conclusion generally drawn from these results is that visuoproprioceptive recalibration (or visuomotor adaptation) relies on temporarily adjusting the weighting of conflicting visual and proprioceptive information to enable adaptive action under specific prior beliefs about one's 'body model'.
The above findings and their interpretation can in principle be accommodated within a hierarchical predictive coding formulation of active inference as a form of Bayes-optimal motor control, in which proprioceptive as well as visual prediction errors can update higher-level beliefs about the state of the body and thus influence action. Hierarchical predictive coding rests on a probabilistic mapping from unobservable causes (hidden states) to observable consequences (sensory states), as described by a hierarchical generative model, where each level of the model encodes conditional expectations ('beliefs') about states of the world that best explains states of affairs encoded at lower levels (i.e., sensory input). The causes of sensations are inferred via model inversion, where the model's beliefs are updated to accommodate or 'explain away' ascending prediction error (a.k.a. Bayesian filtering or predictive coding). Active inference extends hierarchical predictive coding from the sensory to the motor domain, in that the agent is now able to fulfil its model predictions via action. In brief, movement occurs because high-level multi- or amodal beliefs about state transitions predict proprioceptive and exteroceptive (visual) states that would ensue if a particular movement (e.g. a grasp) was performed. Prediction error is then suppressed throughout the motor hierarchy, ultimately by spinal reflex arcs that enact the predicted movement. This also implicitly minimizes exteroceptive prediction error; e.g. the predicted visual action consequences. Crucially, all ascending prediction errors are precision-weighted based on model predictions (where precision corresponds to the inverse variance), so that a prediction error that is expected to be more precise has a stronger impact on belief updating. The 'top-down' affordance of precision has been associated with attention.
This suggests a fundamental implication of attention for behaviour, as action should be more strongly informed by prediction errors 'selected' by attention. In other words, the impact of visual or proprioceptive prediction errors on multisensory beliefs driving action should not only depend on factors like sensory noise, but may also be regulated via the 'top-down' affordance of precision; i.e., by directing the focus of selective attention towards one or the other modality.
Here, we used a predictive coding scheme to test this assumption. We simulated behaviour (i.e., prototypical grasping movements) under active inference, in a simple hand-target phase matching task during which conflicting visual or proprioceptive cues had to be prioritized. Crucially, we included a condition in which proprioception had to be adjusted to maintain visual task performance and a converse condition, in which proprioceptive task performance had to be maintained in the face of conflicting visual information. This enabled us to address the effects reported in the visuomotor adaptation studies reviewed above and studies showing automatic biasing of one's own movement execution by incongruent action observation. In our simulations, we asked whether changing the relative precision afforded to vision versus proprioception-corresponding to attention-would improve task performance (i.e., target matching with the respective instructed modality, vision or proprioception) in each case. We implemented this 'attentional' manipulation by adjusting the inferred precision of each modality, thus changing the degree with which the respective prediction errors drove model updating and action. We then compared the results of our simulation with the actual behaviour and subjective ratings of attentional focus of healthy participants performing the same task in a virtual reality environment. We anticipated that participants, in order to comply with the task instructions, would adopt an 'attentional set' prioritizing the respective instructed target tracking modality over the task-irrelevant one, by means of internal precision adjustment-as evaluated in our simulations.
Results
Results
Simulation results. We based our simulations on predictive coding formulations of active inference. In brief (please see Methods for details), we simulated a simple agent that entertained a generative model of its environment (i.e., the task environment and its hand), while receiving visual and proprioceptive cues about hand posture (and the target). Crucially, the agent could act on the environment (i.e., move its hand), and thus was engaged in active inference.
The simulated agent had to match the phasic size change of a central fixation dot (target) with the grasping movements of the unseen real hand (proprioceptive hand information) or the seen virtual hand (visual hand information). Under visuo-proprioceptive conflict (i.e., a phase shift between virtual and real hand movements introduced via temporal delay), only one of the hands could be matched to the target's oscillatory phase (see Fig. one for a detailed task description). The aim of our simulations was to test whether-in the above manual phase matching task under perceived visuo-proprioceptive conflicts-increasing the expected precision of sensory prediction errors from the instructed modality (vision or proprioception) would improve performance, whereas increasing the precision of prediction errors from the 'distractor' modality would subvert performance. Such a result would demonstrate that-in an active inference scheme-behaviour under visuo-proprioceptive conflict can be augmented via top-down precision control; i.e., selective attention. In our predictive coding-based simulations, we were able to test this hypothesis by changing the precision afforded to prediction error signals-related to visual and proprioceptive cues about hand posture-in the agent's generative model.
Figures two and three show the results of these simulations, in which the 'active inference' agent performed the target matching task under the two kinds of instruction (virtual hand or real hand task; i.e., the agent had a strong prior belief that the visual or proprioceptive hand posture would track the target's oscillatory size change) under congruent or incongruent visuo-proprioceptive mappings (i.e., where incongruence was realized by temporally delaying the virtual hand's movements with respect to the real hand). In this setup, the virtual hand corresponds to hidden states generating visual input, while the real hand generates proprioceptive input.
Under congruent mapping (i.e., in the absence of visuo-proprioceptive conflict) the simulated agent showed near perfect tracking performance (Fig. two). We next simulated an agent performing the task under incongruent mapping, while equipped with the prior belief that its seen and felt hand postures were in fact unrelated, i.e., never matched. Not surprisingly, the agent easily followed the task instructions and again showed near perfect tracking with vision or proprioception, under incongruence (Fig. two). However, as noted above, it is reasonable to assume that human participants would have the strong prior belief-based upon life-long learning and association-that their manual actions generated matching seen and felt postures (i.e., a prior belief that modality specific sensory consequences have a common cause). Our study design assumed that this association would be very hard to update, and that consequently performance could only be altered via adjusting expected precision of vision vs proprioception (see Methods).
Therefore, we next simulated the behaviour (during the incongruent tasks) of an agent embodying a prior belief that visual and proprioceptive cues about hand state were in fact congruent. As shown in Fig. three a, this introduced notable inconsistencies between the agent's model predictions and the true states of vision and proprioception, resulting in elevated prediction error signals. The agent was still able to follow the task instructions, i.e., to keep the (instructed) virtual or real hand more closely matched to the target's oscillatory phase, but showed a drop in performance compared with the 'idealized' agent (cf. Fig. two).
We then simulated the effect of our experimental manipulation, i.e., of increasing precision of sensory prediction errors from the respective task-relevant (constituting increased attention) or task-irrelevant (constituting increased distraction) modality on task performance. We expected this manipulation to affect behaviour; namely, by how strongly the respective prediction errors would impact model belief updating and subsequent performance (i.e., action).
The results of these simulations (Fig. three a) showed that increasing the precision of vision or proprioception-the respective instructed tracking modality-resulted in reduced visual or proprioceptive prediction errors. This can be explained by the fact that these 'attended' prediction errors were now more strongly accommodated by model belief updating (about hand posture). Conversely, one can see a complementary increase of prediction errors from the 'unattended' modality. The key result, however, was that the above 'attentional' alterations substantially influenced hand-target phase matching performance (Fig. three b). Thus, increasing the precision of the instructed task-relevant sensory modality's prediction errors led to improved target tracking (i.e. a reduced phase shift of the instructed modality's grasping movements from the target's phase). In other words, if the agent attended to the instructed visual (or proprioceptive) cues more strongly, its movements were driven more strongly by vision (or proprioception)-which helped it to track the target's oscillatory phase with the respective modality's grasping movements. Conversely, increasing the precision of the 'irrelevant' (not instructed) modality in each case impaired tracking performance.
The simulations also showed that the amount of action itself was comparable across conditions (blue plots in Figs. two and three; i.e., movement of the hand around the mean stationary value of zero point zero five), which means that the kinetics of the hand movement per se were not biased by attention. Action was particularly evident in the initiation phase of the movement and after reversal of movement direction (open-to-close). At the point of reversal of movement direction, conversely, there was a moment of stagnation; i.e., changes in hand state were temporarily suspended (with action nearly returning to zero). In our simulated agent, this briefly increased uncertainty about hand state (i.e., which direction the hand was moving), resulting in a slight lag before the agent picked up its movement again, which one can see reflected by a small 'bump' in the true hand states (Figs. two and three). These effects were somewhat more pronounced during movement under visuo-proprioceptive incongruence and prior belief in congruence-which indicates that the fluency of action depended on sensory uncertainty.
In sum, these results show that the attentional effects of the sort we hoped to see can be recovered using a simple active inference scheme; in that precision control determined the influence of separate sensory modalities-each of which was generated by the same cause, i.e., the same hand-on behaviour by biasing action towards cues from that modality.
Empirical results. Participants practiced and performed the same task as in the simulations (please see Methods for details). We first analysed the post-experiment questionnaire ratings of our participants (Fig. four) to the following two questions: "How difficult did you find the task to perform in the following conditions?" (Q one, answered on a seven-point visual analogue scale from "very easy" to "very difficult") and "On which hand did you focus your attention while performing the task?" (Q two, answered on a seven-point visual analogue scale from "I focused on my real hand" to "I focused on the virtual hand"). For the ratings of Q one, a Friedman's test revealed a significant difference between conditions chi squared sub left parenthesis three point six nine right parenthesis equals forty-seven point one nine, P is less than zero point zero zero one. Post-hoc comparisons using Wilcoxon's signed rank test showed that, as expected, participants reported finding both tasks more difficult under visuo-proprioceptive incongruence (VH incong greater than VH cong, Z sub left parenthesis twenty-three right parenthesis equals four point one four, P is less than zero point zero zero one; RH incong greater than RH cong, Z sub left parenthesis twenty-three right parenthesis equals three point one three, P is less than zero point zero one). There was no significant difference in reported difficulty between VH cong and RH cong, but the VH incong condition was perceived as significantly more difficult than the RH incong condition Z sub left parenthesis twenty-three right parenthesis equals two point five two, P is less than zero point zero five. These results suggest that, per default, the virtual hand and the real hand instructions were perceived as equally difficult to comply with, and that in both cases the added incongruence increased task difficulty-more strongly so when (artificially shifted) vision needed to be aligned with the target's phase.
For the ratings of Q two, a Friedman's test revealed a significant difference between conditions chi squared sub left parenthesis three point six nine right parenthesis equals thirty-five point eight three, P is less than zero point zero zero one. Post-hoc comparisons using Wilcoxon's signed rank test showed that, as expected, participants focused more strongly on the virtual hand during the virtual hand task and more strongly on the real hand during the real hand task. This was the case for congruent (VH cong greater than RH cong, Z sub left parenthesis twenty-three right parenthesis equals three point six five, P is less than zero point zero zero one) and incon- gruent (VH incong greater than RH incong, Z sub left parenthesis twenty-three right parenthesis equals four point zero three, P is less than zero point zero zero one) movement trials. There were no significant differences between VH cong versus VH incong, and RH cong versus RH incong, respectively. These results show that participants focused their attention on the instructed target modality, irrespective of whether the current movement block was congruent or incongruent. This supports our assumption that participants would adopt a specific attentional set to prioritize the instructed target modality.
Next, we analysed the task performance of our participants; i.e., how well the virtual (or real) hand's grasping movements were phase-matched to the target's oscillation (i.e., the fixation dot's size change) in each condition. Note that under incongruence, better target phase-matching with the virtual hand implies a worse alignment of the real hand's phase with the target, and vice versa. We expected (cf. Figure one; confirmed by the simulation results, Figures two and three) an interaction between task and congruence: participants should show a better target phase-matching of the virtual hand under visuo-proprioceptive incongruence, if the virtual hand was the instructed target modality (but no such difference should be significant in the congruent movement trials, since virtual and real hand movements were identical in these trials). All of our participants were well trained, therefore our task focused on average performance benefits from attention (rather than learning or adaptation effects).
The participants' average tracking performance is shown in Figure five. A repeated-measures ANOVA on virtual hand-target phase-matching revealed significant main effects of task F left parenthesis one comma twenty-two right parenthesis equals thirty-one point six nine, P is less than zero point zero zero one and congruence F left parenthesis one comma twenty-two right parenthesis equals one seventy-three point four two, P is less than zero point zero zero one and, more importantly, a significant interaction between task and congruence F left parenthesis one comma twenty-two right parenthesis equals fifty point six nine, P is less than zero point zero zero one. Post-hoc t-tests confirmed that there was no significant difference between the VH cong and RH cong conditions t left parenthesis twenty-three right parenthesis equals one point one nine, P equals zero point two five, but a significant difference between the VH incong and RH incong conditions t left parenthesis twenty-three right parenthesis equals six point five nine, P is less than zero point zero zero one. In other words, in incongruent conditions participants aligned the phase of the virtual hand's movements significantly better with the dot's phasic size change when given the 'virtual hand' than the 'real hand' instruction. Furthermore, while the phase shift of the real hand's movements was larger during VH incong greater than VH cong t left parenthesis twenty-three right parenthesis equals nine point three seven, P less than zero point zero zero one-corresponding to the smaller phase shift, and therefore better target phase-matching, of the virtual hand in these conditions-participants also exhibited a significantly larger shift of their real hand's movements during RH incong greater than RH cong t left parenthesis twenty-three right parenthesis equals four point three one, P is less than zero point zero zero one. Together, these results show that participants allocated their attentional resources to the respective instructed modality (vision or proprioception), and that this was accompanied by significantly better target tracking in each case-as expected based on the active inference formulation, and as suggested by the simulation results.