NNEWSLIVE
HomeBusinessObject groups mutually bias each other’s perceived average size, as evidenced by neural and behavioural data
Business

Object groups mutually bias each other’s perceived average size, as evidenced by neural and behavioural data

Introduction The human visual system demonstrates remarkable parallel processing by rapidly forming ensemble representations, efficiently summarising or averaging object features across groups - such as size 1 , distance 2 , and orientation

E
Editorial Team
October 11, 2026
60 min read
Introduction The human visual system demonstrates remarkable parallel processing by rapidly forming ensemble representations, efficiently summarising or averaging object features across groups - such as size 1 , distance 2 , and orientation 3 . For example, size averaging enables observers to rapidly (in less than 50 ms) and accurately estimate the average size of a set of stimuli, even when they are unable to report information about individual items 1 , 4 , 5 , 6 . Notably, observers can accurately report the average size of a group of objects, even when the visibility of some objects is impaired, as for instance, by object-substitution masking (OSM) 7 . Size averaging performance remains resilient to changes in object features such as set size and set density 8 , and summary statistics are calculated simultaneously for multiple groups of objects without any drop in performance compared to sequential presentation. Overall, these findings indicate that the respective computations occur in parallel and at early stages of the visual processing hierarchy 9 . EEG evidence indicates that ensemble representations, such as average size, are formed even before individual objects are processed 10 . Although these findings suggest that summary statistics are generated early on, the fact that summary statistics are calculated for different groups of objects in parallel 8 implies that preattentive grouping processes must occur first. Taken together, these findings support the idea that individual objects and ensemble representations of objects are represented differently in the visual system, and that information about ensemble representations for different object groups is available relatively early and simultaneously across groups. It is yet unclear whether such different summary statistics constitute independent descriptors, whether they are integrated at some level of processing, and whether such an interaction becomes perceptually relevant. In a recent study, we found evidence that summary size statistics may not be coded independently but instead influence each other in a contrast-like fashion 11 . We observed that the perceived average size of a set of objects was altered by stimuli that induced the Ebbinghaus illusion (Ebbinghaus, 1902). Large Ebbinghaus inducers led to a decrease, whereas small inducers led to an increase in the perceived average size of the target set, indicating a contrast-driven modulation instead. The pattern is consistent with the idea that one group of objects acts as the context for another set and that these contextual objects serve as benchmarks for size judgments, thereby producing a contrast-like effect 12 . This contrast effect may arise from local feature interactions within individual objects or from higher-level operations over summary statistics. In Memis et al.‘s experiment 11 , the sizes of Ebbinghaus inducers were homogeneous within each trial (either large or small) so that their average size and the size of a single inducer object were identical. This confound makes it difficult to uniquely determine whether the observed contrast effects were driven by the overall summary size statistics or by the specific features of the individual inducers. Therefore, in the present experiment, we used heterogeneous sets of stimuli with variable sizes while keeping their average size constant to disentangle individual object size from group average size. In particular, we employed behavioural measures and fMRI to investigate whether the perceived average size of one ensemble influences the perceived average size of another ensemble in a contrast-like manner and whether this contrast effect alters stimulus coding in early visual regions. Participants were presented with two colour-defined ensembles and asked to report the perceived average size of the target ensemble. Target and distractor ensembles were presented in different quadrants of the visual field, allowing separation of their neural signatures in early retinotopic regions of the visual cortex. Furthermore, spatially separating the different stimulus groups enabled us to quantify the effects of mutual size contrasts on the early coding of objects. In addition, based on Murray et al.‘s 13 finding that increased perceived size correlates with larger representations in early visual regions, we hypothesised that: (1) neural activation in regions encoding the target ensemble would be larger when surrounded by a small distractor ensemble as compared to a large ensemble; (2) neural activation in regions encoding the distractor ensemble would be larger when surrounded by a small target ensemble compared to a large target ensemble; and (3) these neural patterns would correspond with behavioural results showing a size-contrast effect. Such mutual influence in the perceived average sizes of different ensembles would suggest that size-contrast effects, as observed in size-contrast illusions, occur at the level of statistical summary descriptors. Materials and methods Participants Twenty-nine healthy participants ( M = 29.28 years, SD = 4.54, 11 females) took part in the fMRI experiment. All participants had normal or corrected-to-normal vision and had normal colour vision as tested with the Velhagen-Broschmann pseudoisochromatic plates 14 . Participants provided written informed consent before the experiment. They were remunerated for their time with 20 euros per hour for the fMRI session and 15 euros per hour for the behavioural session. The study was approved by the ethics committee of the German Society of Psychology (Application ID: WeidnerRalph2024-07-31-VA). All methods were performed in accordance with the Declaration of Helsinki. We determined the sample size based on the effect size reported in Experiment 4 A of our previous behavioural study 11 , where a comparable size-related effect was observed. In that study, the effect was large ( η 2 p = 0.702), corresponding to Cohen’s f = 1.53. For the power analysis, we used a more conservative estimate of f = 0.40. A G*Power analysis for a repeated-measures ANOVA, with 95% power and an alpha level of 0.05, indicated a minimum sample size of 15 participants 15 . As this estimate was based on behavioural data and the expected effect size for the neural measures was unknown, we aimed to recruit 30 participants, corresponding to about twice the minimum sample size indicated by the power analysis, and we successfully tested 29 participants in total. Stimuli The experiment was run via Visual Studio Code 1.68.1 using PsychoPy 2021.2.3 scripts 16 . The stimuli were presented within a 40° visual angle grey background (49.90 cd/m2), with the remainder of the screen being black. The experimental stimuli consisted of two uniformly coloured ensembles (eighteen green and eighteen red circles). During the task, one colour-defined ensemble (e.g., green) was designated as the target ensemble, whereas the other colour-defined ensemble (e.g., red) served as the distractor ensemble and remained task-irrelevant. Due to the sample size, full counterbalancing of colour assignment across target and distractor ensembles was not feasible: for 15 participants, the green ensemble was assigned as the target ensemble, while for 14 participants, the red ensemble was assigned as the target ensemble. The luminance values were measured on the computer used to run the experiment, with green circles at 76.71 cd/m 2 and red circles at 18.27 cd/m 2 . In each trial, one quadrant of the display was randomly assigned as the target quadrant and included fourteen target circles (Fig. 1 , Time Window 2). The quadrant diagonal to the target quadrant was assigned as the distractor quadrant, containing six distractor circles. The remaining two adjacent quadrants included six distractor circles and two additional target circles. This way, participants had to attend to the whole screen to perform the task. The average size of the two additional target circles in the adjacent quadrants was equal to the average size of the target circles on that trial. Hence, the overall average size of the target ensemble remained unchanged. The average sizes of both ensembles were either small (0.7° visual angle) or large (1.3° visual angle), resulting in four experimental conditions: (1) small-average-size target ensemble / large-average-size distractor ensemble (Fig. 1 A), (2) small-average-size target ensemble / small-average-size distractor ensemble (Fig. 1 B), (3) large-average-size target ensemble / small-average-size distractor ensemble (Fig. 1 C), and (4) large-average-size target ensemble / large-average-size distractor ensemble (Fig. 1 D). For clarity, conditions (1) and (3) are referred to as negative and positive size-contrast, respectively, corresponding to the expected underestimation and overestimation of the average size of the target ensemble, whereas conditions (2) and (4) are referred to as size-match. The experimental conditions were presented block-wise, with each block containing trials from only one condition. The individual circle sizes in target and distractor ensembles were determined using a normal cumulative distribution. Specifically, the size of each circle was varied with a standard deviation of 0.15° visual angle around the corresponding mean (0.7° or 1.3° visual angle), and a minimum edge-to-edge distance of 0.5° visual angle was maintained between adjacent circles. The distance from the fixation cross to the edge of the closest circle was approximately 2.65° visual angle for the small ensemble and 2.35° visual angle for the large ensemble. In this way, we manipulated the heterogeneity of the ensembles in each trial while keeping the spatial distribution comparable across trials and across different averages of the target and distractor ensembles. We used the method of constant stimuli to detect the perceived average size of the target ensemble. The size of the comparison circle varied around the average size of the target ensemble in 0.1° increments, resulting in two different comparison-size lists for small (0.7°) and large (1.3°) averages. Ten different comparison sizes were used (half were smaller than the average size of the target ensemble, and the other half were larger). Fig. 1 Illustration of the main task in which the target ensemble consisted of green circles and the distractor ensemble consisted of red circles. Panels ( A - D ) correspond to the four experimental conditions. In the size-match conditions, the average sizes of both ensembles were identical, either small ( B ) or large ( D ). In the negative size-contrast condition ( A ), the average size of the target ensemble was smaller than that of the distractor ensemble. Conversely, in the positive size-contrast condition ( C ), the average size of the target ensemble was larger than that of the distractor ensemble. Procedure Each trial started with a fixation period in which only a fixation cross was presented over a grey background for 1000 ms (Fig. 1 ). In the fMRI session (not in the behavioural pre-test), one of five fixation durations (500 ms, 1000 ms, 1500 ms, 2000 ms, and 2500 ms) was randomly assigned at the beginning of each trial to introduce temporal jitter (Fig. 1 , Time Window 1). Following this, target and distractor ensembles appeared in their respective quadrants for 32 ms. After a 1000 ms fixation period, a comparison circle was displayed at the centre of the screen for 1300 ms. If participants did not respond within the 1300 ms time frame, the trial was considered incorrect. Each trial ended with a 500 ms grey noise pattern. Participants completed 120 trials per block, resulting in a total of 480 trials (2 target ensemble average sizes × 2 distractor ensemble average sizes × 10 comparison sizes × 4 quadrants × 3 repetitions), which lasted approximately 37 min. Participants were instructed to indicate whether the comparison circle was larger or smaller than the average size of the target ensemble (Fig. 1 , Time Window 4). Response mapping was counterbalanced across trials: in half of the trials, a left button press indicated that the comparison circle was larger, whereas in the remaining trials, it indicated that the comparison circle was smaller. The response mapping was indicated during the response period (e.g., larger/left, smaller/right). Participants attended two sessions: a pre-test session and an MRI session, conducted on different days. Initially, participants attended a pre-test session lasting approximately 1 h, during which they completed a colour vision test 14 , an Ebbinghaus screening task (see the Ebbinghaus Screening section), and a practice run of the main task. Participants attended an MRI session 1–4 days after the pre-test session. For four participants, the time between the pre-test and the MRI session was longer (ranging from 8 to 27 days). Those participants received an additional 5 min of practice before scanning. The MRI session lasted about 1.5 h and included a position-localizer task and the main task performed inside the scanner. Ebbinghaus screening During the pre-test session, participants completed a size-judgment task designed to assess the strength of the Ebbinghaus illusion and to test whether a potential size-contrast effect in the main experiment is correlated with an individual’s susceptibility to the Ebbinghaus illusion. The target circle (0.9° visual angle) was presented simultaneously with either small (0.7° visual angle) or large inducers (1.3° visual angle) (Fig. 2 ). Target and inducers appeared randomly at one of the four corners of an imaginary square (10° visual angle) centred on the fixation cross. The same luminance values were used in the main experiment. The perceived size of the target was measured using the method of constant stimuli. The size of the comparison circle varied around the target size (0.9° visual angle) in 0.1° increments. In total, ten different comparison sizes were used (half were smaller than the target size, and the other half were larger). Participants were instructed to maintain central fixation and to ignore the inducers. The task was to indicate whether the comparison circle (Fig. 2 , Time window 4) was seen as larger or smaller than the target circle (Fig. 2 , Time window 2). Each trial started with a 1000 ms fixation period, followed by a 32 ms presentation of the target circle surrounded by either small or large inducers to manipulate the perceived size of the target circle (Fig. 2 A-B). After a 320 ms fixation period, a comparison circle appeared at the centre of the screen, and participants indicated whether it was perceived larger or smaller than the target circle. Participants pressed the right key if the comparison circle was larger and the left key if it was smaller. Each trial ended with a 500 ms screen containing a visual noise pattern. Participants completed 80 trials (2 inducer sizes x 10 comparison sizes x 4 repetitions), which took approximately 5 min. Fig. 2 Illustration of the task in Ebbinghaus screening. Panels ( A ) and ( B ) show conditions with large and small inducers, respectively. The green circle represents the target (Time Window 2), and the red circles represent the inducers (Time Window 2). Eye-tracking data acquisition Eye movement data were recorded from the right eye during the pre-test session using an Eyelink 1000 (SR Research, Mississauga, Ontario, Canada) at a sampling rate of 500 Hz. Participants performed a five-point calibration and validation procedure. A circular area with a radius of 1.5° visual angle around the fixation cross was defined as a region of interest (ROI). We recorded eye movement data throughout the entire experiment, but analysed only critical periods (Fig. 1 , Time Window 1-2-3). The average coordinates of the fixation cross for each trial were employed as a drift check for that specific trial. Preprocessing of the eye movement data was performed using RStudio Version 4.2.0 17 . Behavioural data analysis Perceived object sizes were calculated and analysed across all experimental sessions. To quantify participants’ perception of average size, psychometric curves were generated for each condition by analysing response proportions at 0.1° intervals, reflecting the likelihood of judging the comparison circle as larger than the target circle. We used a logistic function to model this probability ( P ), and the Point of Subjective Equality (PSE) was calculated as P = .5, reflecting the size at which the comparison circle was perceived as equal in size to the target circle. A 2 × 2 repeated-measures ANOVA was conducted on PSEs to examine the effects of the average size of the target and distractor ensembles (small vs. large) in the pre-test and fMRI session. When fitting the psychometric curves to the data, we calculated the goodness-of-fit value, defined as 1 minus the ratio of the residual variance to the total variance in response proportions. The obtained curves demonstrated strong fits in the Ebbinghaus screening task ( r ranged from 0.770 to 0.998), the pre-test ( r ranged from 0.882 to 0.997), and the fMRI session ( r ranged from 0.886 to 0.994). To investigate the relationship between participants’ susceptibility to the Ebbinghaus illusion and size-contrast effects measured in both the pre-test and fMRI task, we calculated individual size-contrast effects. We examined their correlations with the Ebbinghaus illusion strength. We computed size-contrast effects separately for negative and positive size-contrast conditions. Negative size-contrast refers to a condition in which contextual contrast leads to a decrease in perceived size, whereas positive size-contrast indicates an increase in perceived size. These two variants were calculated by using the formulas: $$\:Negative\:Size\_Contrast\:Effect=$$ $$\:PSE\left(size\_match\:condition\right)-PSE\left(negative\:size\_contrast\:condition\right)$$ $$\:Positive\:Size\_Contrast\:Effect=$$ $$\:PSE\left(positive\:size\_contrast\:condition\right)-PSE\left(size\_match\:condition\right)$$ The strength of the Ebbinghaus illusion was quantified as the percentage difference between PSE values obtained with small and large inducers, relative to the target size, using the formula: $$\:Illusion\:strength\:\left(\%\right)\:=\:\frac{(PSE\:small\:inducer-\:PSE\:large\:inducer)\:\times\:\:100}{target\:size=0.9}$$ fMRI measurement Data acquisition Each scanning session included both structural (T1-weighted; TR = 2500 ms, TE = 2.22 ms, flip angle = 7°, FOV = 240 × 240 mm, voxel size = 0.94 × 0.94 × 0.94 mm, number of slices = 208, slice thickness = 0.94 mm) and functional (T2-weighted multiband gradient-echo EPI sequence; TR = 800 ms, TE = 37 ms, flip angle = 52°, FOV = 280 × 280 mm, voxel size = 2 × 2 × 2 mm, number of slices = 72, slice thickness = 2 mm, multiband acceleration factor = 8, phase-encoding acceleration factor = 3) MRI acquisitions. Functional magnetic resonance imaging (fMRI) data were acquired to measure blood oxygenation level-dependent (BOLD) signal changes during task performance. Functional scans were acquired using a 3-T PRISMA MRI system (Siemens, Erlangen, Germany). During imaging, visual stimuli were presented via binocular video goggles (NNL, Bergen, Norway) attached to the 64-channel head coil and adjusted to fit the participants’ vision. Participants were provided with two 5-button response units (Psychology Software Tools Celeritas, Sharpsburg, PA, USA) and used their index fingers to indicate responses. The button on the left hand corresponded to a “left” response, while the button on the right hand corresponded to a “right” response. In the position localizer task, 450 volumes were acquired. In the main task, 2790 volumes were obtained. Data preprocessing The fMRI data were analysed using the statistical parametric mapping software SPM25 (Wellcome Department of Imaging Neuroscience, London; http://fil.ion.ucl.ac.uk/spm/software/spm25 ). Functional images from both the main experiment and the position localizer task were first realigned to correct for inter-scan movement by aligning each image to the participant’s mean functional image. Next, each participant’s mean image was normalised to the standard MNI single-subject template using the unified segmentation approach in SPM25. Finally, to improve the signal-to-noise ratio and compensate for minor anatomical differences while preserving spatial precision in the early visual cortex, the images were smoothed using a 2-mm full-width at half-maximum (FWHM) Gaussian kernel. fMRI analysis From an initial sample of 29 participants, 7 were excluded from all subsequent analyses due to poor task performance and head motion. One participant was excluded due to a high number of missed responses during the task (more than 10% of task responses, exceeding two standard deviations above the mean across participants). The remaining six participants were excluded due to head motion during scanning. Head motion was quantified using frame-wise displacement (FD), calculated as the sum of absolute differences in translational and rotational motion parameters between consecutive volumes 18 , 19 . To ensure data quality, a mean FD threshold of 0.2 mm was applied. The final sample consisted of 22 participants ( M = 28.55 years, SD = 3.91, 8 females, 14 males) and was included in all behavioural and functional analyses. Position localizer task Before the fMRI experiment, a position-localizer task was conducted to identify cortical representations in each quadrant of the visual field (lower-left, lower-right, upper-left, and upper-right). This task involved presenting a contrast-reversing flickering checkerboard featuring black and white squares in each quadrant at 8 Hz. Each checkerboard covered a 20° visual angle and consisted of a 10 × 10 grid (100 squares), with each square subtending a 2° visual angle. Each checkerboard was presented for 18 s. The entire task lasted approximately 6 min. Four regressors indicated the onsets of 18-seconds visual stimulations, with each regressor corresponding to a different quadrant. The hemodynamic response for each condition was modelled using a canonical HRF and its time derivative, with head movement parameters included as additional regressors in the design matrix. For each participant, first-level analyses were performed to test for larger BOLD amplitudes in one condition relative to the remaining three, yielding condition-specific differential contrasts. These subject-level contrasts were thresholded at p < .001 (whole-brain FWE corrected at the peak voxel level) to define functional masks representing retinotopically distinct stimulus locations for each quadrant (Fig. 6 ). Next, each participant’s quadrant-specific functional mask from the localizer was intersected with probabilistic maps of early visual areas (V1v, V1d, V2v, V2d, V3v, V3d) provided by Wang et al. 20 , registered in MNI space. This procedure resulted in subject-specific functional ROIs defined by both visual area (e.g., V1) and quadrant representation (e.g., upper-left). These subject-specific functional ROIs were then used in the main task to quantify the number of significantly active voxels in the experimental conditions. Main task During the scanning session, participants performed the same size-averaging task as in the pre-test session, in which they indicated whether the comparison circle was larger or smaller than the average size of the target ensemble, while ignoring a simultaneously presented distractor ensemble (Fig. 1 ). Initially, we defined 16 onset regressors, corresponding to the 16 experimental conditions (2 target ensemble average sizes: small, large; x 2 distractor ensemble average sizes: small, large; x 4: quadrants: lower-left, lower-right, upper-left, upper-right). This way, we could test the effects of our experimental manipulations separately in each quadrant. The hemodynamic response was modelled using a canonical HRF and its time derivative, with head movement parameters included as additional regressors in the design matrix. ROI-analyses: Separating target and distractor ensembles To investigate the mutual size-contrast effect between target and distractor ensembles, we first needed to disentangle their neural signals. We defined subject-specific functional ROIs based on the localizer scans to test functional activation in voxels representing each quadrant (lower-left, lower-right, upper-left, upper-right) in turn. This allowed us to examine the neural representation of the target ensemble in isolation, as well as the influence of the distractor ensemble in a different quadrant (diagonally opposite the target ensemble). Later, the quadrant-specific functional ROIs were further subdivided into different visual areas (V1, V2, V3) for each participant. Testing for mutual influences: Size-contrast vs. Size-match To test mutual size‐contrast effects across ensembles, we compared the number of suprathreshold voxels for an ensemble of interest (e.g., target ensemble with a large average size) as a function of the average size of a context ensemble (e.g., distractor ensemble with a small average size). The average size of the context ensemble could either match the average size of the ensemble of interest, constituting a size-match condition (e.g., target ensemble with a large average size in combination with a distractor ensemble with a large average size), or it could differ from the average size of the ensemble of interest, constituting a size-contrast condition (e.g., target ensemble with a large average size in combination with a distractor ensemble with a small average size). If the average size of the context ensemble affects the neural coding of the ensemble of interest, we would expect the number of voxels activated by the ensemble of interest to differ between the size-contrast and size-match conditions. This effect can be quantified as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left(size\_contrast\right)-{N}_{vox}(size\_match)$$ where: \(\:{N}_{vox}(.)\) =number of suprathreshold voxels in the condition indicated in parentheses; “size‐contrast” = an ensemble of interest paired with a context ensemble of a different average size; “size‐match” = an ensemble of interest paired with a context ensemble of the same average size; Positive \(\:\varDelta\:{N}_{vox}\) values indicate more activated voxels in the size-contrast condition and negative values indicate fewer. Based on the size-contrast hypothesis, we can make clear directional predictions about \(\:\varDelta\:{N}_{vox}\) . As indicated below in more detail, a positive size-contrast, in which a large ensemble of interest is presented with a smaller context ensemble, is expected to yield a positive \(\:\varDelta\:{N}_{vox}\) . In contrast, a negative size-contrast, where a small ensemble of interest is presented with a large context ensemble, is hypothesised to generate a negative \(\:\varDelta\:{N}_{vox}\) . Large ensembles of interest Pairing a large ensemble of interest with a small context generates a positive size-contrast and should increase its perceived size, leading to more activated voxels. In that case \(\:\varDelta\:{N}_{vox}\) is expected to be > 0. For large target ensembles, this can be described as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left({T}_{large}|{D}_{small}\right)-\left({T}_{large}|{D}_{large}\right)>0$$ For large distractor ensembles, this can be described as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left({D}_{large}|{T}_{small}\right)-\left({D}_{large}|{T}_{large}\right)>0$$ where for all specific-case formulas: \(\:{N}_{vox}\left(X\right|Y)\) = number of suprathreshold voxels for ensemble X (ensemble of interest) when presented with ensemble Y(context ensemble); T = target ensemble, D = distractor ensemble; Subscript large/small = average‐size category of that ensemble; The vertical bar “|” means “presented with”; \(\:\varDelta\:{N}_{vox}\) =voxel-count difference between size‐contrast and size‐match conditions. Small ensembles of interest Pairing a small ensemble of interest with a large context should result in negative size-contrast and decrease its perceived size, leading to fewer activated voxels. In that case \(\:\varDelta\:{N}_{vox}\) is expected to be < 0. For small target ensembles, this can be described as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left({T}_{small}|{D}_{large}\right)-\left({T}_{small}|{D}_{small}\right)<0$$ For small distractor ensembles, this can be described as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left({D}_{small}|{T}_{large}\right)-\left({D}_{small}|{T}_{small}\right)<0$$ To summarise, mutual size‐contrast interactions between different ensembles are reflected in differences between the size‐contrast and size‐match conditions. The resulting voxel‐count difference \(\:\varDelta\:{N}_{vox}\) is expected to be positive for positive size‐contrasts (ensemble of interest larger than context) and negative for negative size‐contrasts (ensemble of interest smaller than context). Functionally, four differential contrasts per quadrant were defined to compare the size-match conditions with the size-contrast conditions (Size-match > negative Size-contrast and positive Size-contrast > Size-match) for target (Fig. 7 A-B) and distractor (Fig. 7 C-D) ensembles. Statistical testing We used voxel count rather than mean BOLD amplitude because size-related perceptual effects are thought to modulate the spatial extent of cortical activation. Voxel count, therefore, provides a sensitive measure of activation spread within retinotopic regions, consistent with the evidence that perceived size is well reflected in differences in cortical extent of activity 13 . Using the functionally defined and intersected subject-level ROIs, we extracted the number of significantly active voxels in each functional ROI from the main task’s differential contrasts, applying a threshold of p < .001 (uncorrected). The resulting voxel count served as the dependent variable. To control for variability in functional ROI size, the extracted voxel counts were normalised by calculating the percentage of activated voxels relative to the total voxel count within each functional ROI, enabling comparisons across visual areas and quadrants. The resulting differences were submitted to second-level one-tailed one-sample t-tests, with Holm correction applied across V1, V2, and V3 separately for each contrast. Because these contrasts have an expected false-positive rate of 0.1% given the uncorrected voxel-wise threshold of p < .001, we did not test the percentage differences in activated voxels against zero. Instead, for each ROI, we compared the observed percentage difference to the percentage of voxels expected to pass the voxel-wise threshold by chance alone (0.1% of voxels within the ROI). Percentages significantly above 0.1% indicate that the functional ROI showed a reliable difference in activation between conditions. In other words, this would be taken as evidence of increased modulation in one condition (e.g., positive size-contrast) relative to the other (e.g., size-match condition). Similarly, greater modulation is expected when comparing the size-match condition with the negative size-contrast. Results Behavioural results Eye movement data Eye movement data were analysed for the twenty-two participants. We employed a 2 × 2 repeated measures ANOVA to examine fixation maintenance within the fixation ROI across experimental conditions, with factors of the average size of the target ensemble (small, large) and the average size of the distractor ensemble (small, large). During the critical periods of the experiment, participants’ gaze was located within the fixation ROI for an average of 96.70% of the time. The ANOVA did not reveal significant main effects of the average size of the target ensemble [small vs. large] ( F (1, 21) = 0.53, p = .474, η 2 p = 0.025), and average size of the distractor ensemble [small vs. large] ( F (1, 21) = 0.34, p = .567, η 2 p = 0.016). No significant interaction was observed between the two factors ( F (1, 21) = 0.22, p = .646, η 2 p = 0.010), indicating comparable fixation maintenance across all experimental conditions. Ebbinghaus screening Figure 3 displays the mean PSEs (A) for the small and large inducer conditions, and psychometric curves for the single-subject (B) and group-level (C). The Ebbinghaus screening task was conducted to assess participants’ susceptibility to size-contrast effects. A two-tailed paired-sample t-test indicated a significant difference in the mean PSE values for small and large inducers ( t (21) = 7.51, p < .001, Cohen’s d = 1.601). Specifically, participants estimated the target size as significantly larger in the small inducer condition ( M = 0.78, SE = 0.03) than in the large inducer condition ( M = 0.66, SE = 0.02), showing that the Ebbinghaus screening task significantly altered the perceived size of the target stimulus. Overall, the illusion strength was approximately 13.77% across participants. Fig. 3 Perceived target size and psychometric curves in Ebbinghaus screening. ( A ) Averaged PSEs across different inducer types were plotted. The blue bar represents the small inducer condition, and the orange bar indicates the large inducer condition. Solid grey lines connect each participant’s performance across inducer types. The horizontal dashed grey line represents the physical size of the target stimulus. Asterisks (*) represent significant differences at p < .05. Single-subject ( B ) and group-level ( C ) psychometric curves show the proportion of “larger” responses as a function of comparison circle size. The horizontal grey line marks the 50% response level, and the vertical grey lines indicate the corresponding PSEs. Error bars indicate the standard errors around the mean for within-subject contrasts 21 . Pre-test Figure 4 A represents the mean PSEs for the average size of the target and distractor ensembles. Figure 4 B and C represent single-subject and group-level psychometric curves, respectively. As expected, participants perceived the average size as significantly smaller in the small-average-size target ensemble condition ( M = 0.59, SE = 0.02) than in the large-average-size target ensemble condition ( M = 1.24, SE = 0.02). Additionally, the estimated perceived size was larger in the small-average-size distractor ensemble condition ( M = 0.97, SE = 0.02) compared to the large-average-size distractor ensemble condition ( M = 0.86, SE = 0.02). A 2 × 2 ANOVA on mean PSEs revealed significant main effects of the average size of the target ensemble [small vs. large] ( F (1, 21) = 1384.40, p < .001, η 2 p = 0.985), and the average size of the distractor ensemble [small vs. large] ( F (1, 21) = 62.88, p < .001, η 2 p = 0.750). However, there was no interaction between the two factors ( F (1, 21) = 0.16, p = .698, η 2 p = 0.007). Given the luminance differences between green and red stimuli, we tested whether luminance influenced perceived size. To do this, we included the colour assigned to the target ensemble (green vs. red) as a between-subjects factor. There was a significant main effect of colour assignment ( F (1, 20) = 4.53, p = .046, η 2 p = 0.185), with the perceived average size of the target ensemble being larger in the red version ( M = 0.95, SE = 0.03) than in the green version ( M = 0.88, SE = 0.03). Importantly, there was no interaction between colour assignment and the average size of the target ensemble ( F (1, 20) = 1.02, p = .324, η 2 p = 0.049), or between colour assignment and the average size of the distractor ensemble ( F (1, 20) = 1.79, p = .196, η 2 p = 0.082). The three-way interaction was also not significant ( F (1, 20) = 0.78, p = .387, η 2 p = 0.038). Thus, the main effect of colour assignment reflected an overall shift in perceived average size of the target ensemble across all experimental conditions, rather than a differential influence on the effects of target or distractor ensemble size. To directly assess the size-contrast effect, two-tailed paired-sample t-tests were conducted comparing each contrast condition to the size-match condition (Fig. 4 ). By design, the size-contrast manipulation was defined in two directions: (1) negative size-contrast, where the average size of the target ensemble was smaller than that of the distractor ensemble (Fig. 1 A), hypothesized to lead to underestimation; and (2) positive size-contrast, where the average size of the target ensemble was larger than that of the distractor ensemble (Fig. 1 C), and was expected to result in overestimation. Overall, the perceived average size of the target ensemble was altered by the average size of the distractor ensemble. For negative size-contrast, participants perceived small-average-size target ensembles as significantly smaller when they were presented alongside large-average-size distractor ensembles ( M = 0.54, SE = 0.02), compared to when they were presented with small-average-size distractor ensembles ( M = 0.64, SE = 0.02), t (21) = 5.06, p < .001, Cohen’s d = 1.078. In contrast, positive size-contrast produced the opposite pattern: participants perceived large-average-size target ensembles as significantly larger when they were presented with small-average-size distractor ensembles ( M = 1.29, SE = 0.03) compared to when they were presented alongside large-average-size distractor ensembles ( M = 1.18, SE = 0.02), t (21) = 5.11, p < .001, Cohen’s d = 1.088. Together, these results demonstrate a robust size-contrast effect. Specifically, the perceived average size of the target ensemble was systematically modulated by the size of the surrounding distractor ensemble, despite this ensemble being task-irrelevant. This finding is consistent with our hypothesis that the average size of a distractor ensemble influences the perceived average size of the target ensemble. Fig. 4 Perceived average size and psychometric curves in the pre-test. ( A ) The averaged PSEs were plotted against the average sizes of the target and distractor ensembles. Bars on the left reflect conditions with small-average-size target ensembles, and bars on the right reflect conditions with large-average-size target ensembles. The blue bars represent conditions with small-average-size distractor ensembles, while the orange bars represent those with large-average-size distractor ensembles. Individual data points are overlaid: open and filled circles represent participants for whom the red and green circles served as the target ensemble, respectively. Grey lines connect each participant’s performance across conditions. The horizontal black lines represent the physical average size of the target ensemble. The figures below the x-axis illustrate the corresponding experimental conditions. Asterisks (*) represent significant differences at p < .05. Single-subject ( B ) and group-level ( C ) psychometric curves show the proportion of “larger” responses as a function of comparison circle size. Solid lines represent size-match conditions (orange: small-average-size target ensemble, blue: large-average-size target ensemble) and dashed lines represent size-contrast conditions (orange: negative, blue: positive). The horizontal grey line marks the 50% response level, and the vertical grey lines indicate the corresponding PSEs. Error bars indicate the standard errors around the mean for within-subject contrasts 21 . fMRI-behaviour Figure 5 shows the mean PSEs (A) for the average sizes of the target and distractor ensembles, and psychometric curves at the single-subject (B) and group-level (C). The behavioural results obtained during the fMRI session replicated the pre-test findings, showing that the perceived average size of the target ensemble was modulated by the size of the distractor ensemble (Fig. 5 ). As in the pre-test, the perceived average was smaller in the small-average-size target ensemble condition ( M = 0.57, SE = 0.02) than in the large-average-size target ensemble condition ( M = 1.23, SE = 0.02). Participants also perceived the average size as larger in the small-average-size distractor ensemble condition ( M = 0.93, SE = 0.02) than in the large-average-size distractor ensemble condition ( M = 0.87, SE = 0.02). The 2 × 2 ANOVA on mean PSEs replicated the pre-test findings, confirming significant main effects of both the average size of the target ensemble [small vs. large] ( F (1, 21) = 1405.43, p < .001, η 2 p = 0.985), and the average size of the distractor ensemble [small vs. large] ( F (1, 21) = 26.52, p < .001, η 2 p = 0.558). The interaction between the two factors was not significant ( F (1, 21) = 0.23, p = .639, η 2 p = 0.011), consistent with the pre-test findings. As in the pre-test, we tested the impact of target luminance on perceived size by including the colour assignment as a between-subjects factor. There was no significant main effect of colour assignment ( F (1, 20) = 1.19, p = .288, η 2 p = 0.056), and no interaction between colour assignment and the average size of the target ensemble ( F (1, 20) = 0.09, p = .765, η 2 p = 0.005), or between colour assignment and the average size of the distractor ensemble ( F (1, 20) = 0.23, p = .635, η 2 p = 0.011). The three-way interaction was also not significant ( F (1, 20) = 0.37, p = .550, η 2 p = 0.018), providing no evidence for either an overall difference in perceived average size of the target ensemble between the colour versions or a differential influence across conditions in this session. Two-tailed paired-sample t-tests confirmed the size-contrast effect, replicating the size-contrast effect in both directions as observed in the pre-test (Fig. 5 A). The negative size-contrast effect resulted in a significant underestimation: participants perceived small-average-size target ensembles as significantly smaller when paired with large-average-size distractor ensembles ( M = 0.54, SE = 0.02) than when paired with small-average-size distractor ensembles ( M = 0.60, SE = 0.02), t (21) = 3.80, p < .001, Cohen’s d = 0.810. The opposite pattern emerged for the positive size-contrast effect: the perceived average size was significantly larger in large-average-size target ensembles when paired with small-average-size distractor ensembles ( M = 1.27, SE = 0.03) than when paired with large-average-size distractor ensembles ( M = 1.20, SE = 0.03), t (21) = 4.50, p < .001, Cohen’s d = 0.960. Fig. 5 Perceived average size and psychometric curves in the fMRI. Plotting conventions and statistical annotations are identical to those in Fig. 4 . ( A ) The averaged PSEs were plotted against the average sizes of the target and distractor ensembles. Single-subject ( B ) and group-level ( C ) psychometric curves show the proportion of “larger” responses as a function of comparison circle size. Correlation analysis Pearson correlation analyses were conducted to examine the relationship between the strength of the Ebbinghaus illusion and the size-contrast effects observed in both the pre-test and the fMRI session. To match the direction of perceptual effects across tasks, PSEs in the Ebbinghaus conditions inducing underestimation (i.e., averaged PSEs from the large inducer condition) and overestimation (i.e., averaged PSEs from the small inducer condition) were correlated with the negative and positive size-contrast effects, respectively. In the pre-test, no significant correlations were observed between PSEs in the underestimation-related Ebbinghaus condition (large inducer condition) and the negative size-contrast effect ( r = − .232, p = .298), nor between PSEs in the overestimation-related Ebbinghaus condition (small inducer condition) and the positive size-contrast effect ( r = − .258, p = .246). The same pattern was observed in the fMRI session: no significant correlations emerged between PSEs in the underestimation-related Ebbinghaus condition and the negative size-contrast effect ( r = .079, p = .728), or between PSEs in the overestimation-related Ebbinghaus condition and the positive size-contrast effect ( r = .119, p = .596). The overall size-contrast effect was also quantified as the difference between the positive and negative size-contrast effects. No significant correlations were observed between the overall Ebbinghaus illusion strength and the overall size-contrast effect in either the pre-test ( r = − .101, p = .656) or the fMRI session ( r = .238, p = .286). Together, these findings suggest that participants’ susceptibility to the Ebbinghaus illusion is unlikely to be associated with size-contrast effects on the level of object groups, implying that size-contrast operates through different mechanisms for ensembles compared to single items, despite producing similar perceptual outcomes. Further, we examined whether the behavioural size-contrast effects were correlated with the corresponding fMRI percentage-difference measures for the target ensemble, as behavioural responses were available only for this ensemble. These correlations were not significant (all p s > 0.05). The presence of group-level effects in both measures, despite the absence of a correlation between them, may indicate that the behavioural and fMRI measures capture distinct aspects of the size-contrast effect, such that individual differences in the behavioural effect are not fully reflected in modulation within retinotopic visual cortex. fMRI results Position localizer Figure 6 displays group-level visualisations of the four quadrant-specific activations derived from the position localizer task. Whole-brain statistical maps were thresholded at p < .001, FWE-corrected at the peak level, to identify significant activation clusters evoked by each quadrant. This analysis revealed distinct peak activations consistent with the retinotopic organisation of the early visual cortex. Specifically, stimuli presented in the lower right quadrant (Fig. 6 , blue) elicited activation in the left hemisphere above the calcarine sulcus, while stimuli in the lower left quadrant (Fig. 6 , green) activated the right hemisphere above the calcarine sulcus. Likewise, stimuli from the upper right quadrant (Fig. 6 , yellow) were represented in the left hemisphere below the calcarine sulcus, and those from the upper left quadrant (Fig. 6 , red) were observed in the right hemisphere below the calcarine sulcus. Notably, activation associated with the lower-right quadrant (Fig. 6 , blue) exhibited a broader lateral spread than those observed for the other quadrants, consistent with previous reports of non-uniform cortical magnification in early visual cortex, including greater cortical representation along the lower relative to the upper vertical meridian 22 . Please note that subject-specific functional ROIs were defined by intersecting each participant’s quadrant-specific localizer mask with the corresponding V1, V2, and V3 masks from the Wang et al. probabilistic atlas 20 . Additional visual areas outside the probabilistic definitions of V1, V2 and V3 were excluded even if they were part of the functional pattern revealed by the corresponding localizer contrast and hence did not influence the main analyses. Fig. 6 Group-level activation patterns from the position localizer task. The left panel shows quadrant-specific retinotopic activations identified through whole-brain analysis, displayed on a 3D surface-rendered brain. The middle panel illustrates the spatial configuration of the stimulus locations used to define retinotopically distinct regions. The right panel presents axial and coronal slices showing the corresponding peak activations, thresholded at p < .001 (FWE-corrected at the peak level). Main experiment Figure 7 illustrates the percentage differences in activated voxels across early visual areas (V1, V2, V3) for four differential contrasts, averaged across visual field quadrants (lower-right, lower-left, upper-right, upper-left). These differential contrasts tested how a context ensemble influenced the neural representation of an ensemble of interest by comparing size-contrast conditions with size-match conditions. The contrasts were calculated separately for large ensembles of interest, testing for positive size-contrast effects ( \(\:\varDelta\:{N}_{vox}>0\) ) and for small ensembles of interest, testing for negative size-contrast effects ( \(\:\varDelta\:{N}_{vox}<0\) ). In addition, differential contrasts were calculated separately for target and distractor ensembles. Target ensemble: small average size We tested whether a negative size-contrast reduced the number of activated voxels. For clarity of presentation, voxel‐count differences were computed as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left(size\_match\right)-{N}_{vox}(size\_contrast)$$ so that the reduced number of activated voxels in the size-contrast condition would generate positive \(\:\varDelta\:{N}_{vox}\) . One-tailed one-sample t-tests against the 0.1% chance level were conducted on the percentage differences in significantly activated voxels extracted from each functional ROI to estimate the negative size-contrast effect (Fig. 7 A). The results revealed that negative size-contrast significantly affected the neural coding of average size in all functional ROIs [V1 ( t (21) = 2.65, p = .021, Cohen’s d = 0.565); V2 ( t (21) = 2.62, p = .021, Cohen’s d = 0.559), and V3 ( t (21) = 2.69, p = .021, Cohen’s d = 0.573)]. Target ensemble: large average size Similarly, we tested for a positive size-contrast effect for target ensembles (Fig. 7 B). Since a positive size-contrast effect is expected to increase the number of activated voxels in the size-contrast condition, the voxel-count differences were computed as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left(size\_contrast\right)-{N}_{vox}(size\_match)$$ One-tailed one-sample t-tests on the resulting voxel counts revealed significantly more activated voxels ( \(\:\varDelta\:{N}_{vox}>0.1\%\) ) in the positive size-contrast condition in all functional ROIs [V1, t (21) = 2.09, p = .024, Cohen’s d = 0.447; V2 ( t (21) = 2.74, p = .012, Cohen’s d = 0.584), and V3 ( t (21) = 2.97, p = .012, Cohen’s d = 0.632)]. Together, these results provide robust evidence that neural populations in early visual areas represent perceived average size differences. In combination with the behavioural evidence of contrast-like modulation in perceived average size, these functional imaging results provide strong support for a size-contrast effect, demonstrating that the perceived average size of the target ensemble was modulated by the size of the distractor ensemble. Distractor ensemble: small average size To explore whether a similar pattern exists for the distractor ensemble, we applied the same analysis for quadrants including distractor ensembles. For a small-average-size distractor ensemble, we tested whether negative size-contrasts reduced the number of activated voxels (Fig. 7 C). Again, for clarity of presentation, voxel‐count differences were computed as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left(size\_match\right)-{N}_{vox}(size\_contrast)$$ so that the reduced number of activated voxels in the size-contrast condition would generate positive \(\:\varDelta\:{N}_{vox}\) . One-tailed one-sample t-tests were conducted on the percentage differences in activated voxels extracted from each functional ROI to estimate the negative size-contrast effect. Similar to the target ensemble analysis, negative size-contrasts ( \(\:\varDelta\:{N}_{vox}>0.1\%\) ) significantly affected the neural representation of the distractor ensemble’s average size across all functional ROIs [V1 t (21) = 2.07, p = .048, Cohen’s d = 0.441; V2 ( t (21) = 2.24, p = .048, Cohen’s d = 0.478), and V3 ( t (21) = 2.30, p = .048, Cohen’s d = 0.491)]. Distractor ensemble: large average size As for target ensembles, we tested whether a positive size-contrast increased the number of activated voxels when contrasting size-contrast and size-match conditions (Fig. 7 D). Positive size-contrast is expected to increase the number of activated voxels in the size-contrast condition; therefore, the voxel-count differences were computed as: $$\:\varDelta\:{N}_{vox}={N}_{vox}\left(size\_contrast\right)-{N}_{vox}(size\_match)$$ The results showed a consistent but non-significant pattern, with more activated voxels across V1 ( t (21) = 1.68, p = .078, Cohen’s d = 0.358), V2 ( t (21) = 1.85, p = .078, Cohen’s d = 0.395) and V3 ( t (21) = 2.21, p = .057, Cohen’s d = 0.472). Overall, although the current behavioural design did not allow direct measurement of perceived size in the task-irrelevant distractor ensemble — since participants were only asked to judge the target — the neural data suggest that distractor objects may also be subject to size-contrast effects: the spatial extent of activation in retinotopic regions representing the distractor ensemble was modulated in a contrast-like manner, with a significant negative size-contrast effect and a corresponding although non-significant pattern for positive size-contrast, paralleling the pattern observed for the target ensemble. Fig. 7 Differences in the percentage of activated voxels across functional ROIs. The figure displays four differential contrasts, with panels ( A ) and ( B ) showing quadrants containing only target ensembles and panels ( C ) and ( D ) showing quadrants containing only distractor ensembles. Each panel illustrates a pairwise comparison between the size-match and the size-contrast condition, followed by the percentage difference in activated voxels within V1 (blue), V2 (red), and V3 (yellow). Individual data points are overlaid: open and filled circles indicate participants for whom the red and green circles served as the target ensemble, respectively. Error bars indicate the standard errors around the mean for within-subject contrasts 21 . Whole-brain covariate analysis Given that individual variability in the behavioural size-contrast effect was not significantly associated with the spatial extent of activation in retinotopic areas, we examined whether this variability was reflected in neural responses elsewhere in the brain. A whole-brain flexible factorial analysis was conducted using first-level contrast images from the positive and negative size-contrast conditions. The behavioural size-contrast effect was entered as a covariate to identify regions in which activation across these conditions varied as a function of the magnitude of this effect. Using a voxel-level cluster-defining threshold of p < .001 (uncorrected) and cluster-level family-wise error (FWE) correction at p < .05, the analysis revealed two significant clusters in the left hemisphere: one in the intraparietal sulcus (IPS; MNI − 40, −52, 42; T = 5.68; k = 55, p = .006) extending toward the angular gyrus, and one in the middle temporal gyrus (MTG; MNI − 66, −36, − 2; T = 4.28; k = 40, p = .034). Taken together with the retinotopic voxel-count results—which showed a perceived-size effect at the condition level but no significant association with the behavioural effect—the whole-brain analysis revealed that the magnitude of the behavioural size-contrast effect instead scaled with activation in higher-level cortical regions. Discussion This study provides strong evidence for a mutual size-contrast effect in ensemble representations, demonstrating that the average size of two groups of objects exerts a mutual influence, with each group acting as context for the other. In particular, behavioural data showed that participants perceived the average size of a target ensemble as larger when presented alongside a smaller distractor ensemble, and as smaller when surrounded by a larger distractor ensemble. This pattern closely replicates classical size-contrast effects, such as those observed in the Ebbinghaus illusion 23 , but illustrates that this phenomenon exists at the level of ensemble representations. Notably, we observed no significant correlation between the participants’ susceptibility to the Ebbinghaus illusion, as detected in the screening task, and the size-contrast effect measured in the main task. This suggests that size-contrast effects may operate via distinct mechanisms for ensemble summary statistics and for individual objects, despite yielding similar perceptual outcomes. Retinotopically defined ROI analysis revealed a neural activation pattern indicating that size‐contrast modulated the activity evoked by a target ensemble in early visual areas, in a manner consistent with perceptual size‐contrast effects. For the distractor ensemble, a similar neural activation pattern was observed, despite the absence of behavioural data confirming a perceived size-contrast effect. The presence of this neural modulation suggests that size-contrast may also have influenced the processing—and potentially the perception—of the distractor ensemble, even though they were not directly attended or relevant to the behavioural task. While earlier studies have shown that summary statistics can be computed simultaneously across object groups 8 , our findings show that these representations are not entirely independent, with size-contrast effects occurring at the level of statistical summary representations. Our neuroimaging results support the involvement of early visual areas in ensemble size computation, rather than being restricted to higher-level statistical operations 24 , 25 . Although previous studies have shown that perceived, not retinal size, is encoded in early visual areas for individual objects 13 , 26 , 27 , 28 , our findings reveal that such a mechanism applies to ensemble representations, with perceived average size affecting retinotopic activation. This account is further supported by the consistency of the behavioural size-contrast effect across individuals (Fig. 4 A), suggesting that the phenomenon reflects a common perceptual process that is shared across observers. Specifically, conditions in which a larger distractor ensemble surrounded a small target ensemble showed fewer activated voxels than the size-match condition, aligning with participants’ behavioural responses, in which the target ensemble appeared smaller. Conversely, conditions involving a large target ensemble surrounded by a smaller distractor ensemble demonstrated an increased number of activated voxels. These results suggest that the visual system dynamically adjusts cortical activation patterns in response to perceived size differences in ensemble representations. The present findings align with and extend previous research showing that the visual system automatically rescales ensemble representations relative to contextual cues before statistical summaries are computed 29 , 30 . Such rescaling implies that ensemble representations reflect perceived rather than retinal sizes of object groups 31 , a claim supported by our functionally defined retinotopic analysis. Furthermore, we replicated and extended our previous findings 11 , in which we initially observed size-contrast effects in which the distractor ensemble influenced the perceived average size of the target ensemble. The current study clarifies these mechanisms, suggesting reciprocal size modulation between the target and distractor ensembles. Our findings indicate that the visual system constructs ensemble representations of object groups regardless of whether they are attended. Notably, even distractor ensembles—those not in the focus of attention—display neural signatures of size-contrast effects, paralleling those seen for attended groups. This implies that ensemble information from both attended and unattended sets is integrated into the scene’s relational framework, influencing how other object groups are perceived and represented. Our results furthermore address the question of whether, in ensemble perception, multiple stimulus sets are processed independently or combined into a single, pooled representation 8 . If all stimuli were pooled into a single representation, one would expect an additive averaging effect, in which the presence of large objects increases the ensemble’s perceived size. Our results argue against this assumption, instead demonstrating a contrast-like interaction between distinct ensembles for object size (but see Ortego & Störmer 32 for different effects for stimulus orientation). An alternative explanation is that the observed size-contrast effect might result from regression to the mean or a global-reference account. In this framework, responses to each set may be shaped not only by subset-level statistics but also by the overall statistics of the display, providing one possible explanation for how the contrast effect might be generated. This interpretation aligns with prior research demonstrating that both subset-level and overall display statistics can influence size representations 8 , 33 . At the neural level, this broader influence of display-level information can be considered in terms of spatial pooling in early visual cortex. Population receptive fields (pRFs) are spatially extended, and centre–surround interactions may allow voxels assigned to one quadrant to receive input from adjacent regions 34 , 35 . Thus, pRF-level pooling and centre–surround interactions may contribute to the observed neural responses. However, if such mechanisms—whether simple additive spillover or biased competition—were responsible, they would predict voxel-count differences in the opposite direction to our findings. Specifically, larger distractors should increase, or at least not reduce, activation in the target-quadrant representation, but our results show the opposite pattern relative to the corresponding size-match conditions. Therefore, spillover from the distractor quadrant is unlikely to be a major driver of the observed voxel-count effects. In principle, such contrast effects may emerge from local changes and the perceived size of individual elements that collectively bias ensemble judgments, rather than originating at the level of ensemble representations 26 , 36 . However, several aspects of the present study make a purely local, element-level account unlikely. Notably, the target and distractor ensembles were presented diagonally, a configuration intended to minimise direct lateral interactions between individual elements from the two sets. Although population receptive fields (pRFs) at quadrant borders may overlap, such overlap would involve only a small subset of elements. The average size of an ensemble, by definition, requires the integration of all its elements across a larger spatial extent — a computation that cannot be performed by border elements alone 1 . Beyond the spatial arrangement, the nature of the dependent variable itself challenges a purely local, element-level explanation. No single V1 neuron or pRF encodes the average size of a group of spatially distributed elements—extracting this parameter requires a pooling process that combines information across a spatial scale larger than individual local representations can provide. Consistent with this, we found no significant correlation between individual participants’ susceptibility to the Ebbinghaus illusion and the magnitude of the ensemble size-contrast effect. If local element-level changes were the primary driver of the effect, such a correlation would be expected, given that the Ebbinghaus illusion is strongly modulated by the spatial configuration and the proximity of inducers to the target 37 . Thus, while local interactions between individual elements may contribute to the effect, the present findings are better captured by an account in which ensemble-level representations participate in contextual size processing. While our findings clearly indicate the involvement of early retinotopic regions, they do not imply that these regions are the sole origin of the observed size-contrast effect. Rather, we identified early retinotopic cortex as a locus where ensemble-level size-contrast is expressed, possibly as part of a larger network involving higher-level regions. Consistent with this, a whole-brain covariate analysis linked individual differences in the behavioural size-contrast effect to size-contrast-related activation in the left IPS and the MTG, whereas no such relationship was observed for retinotopic voxel-count measures. One possibility is therefore that the effect is computed within a broader network of regions, including the left IPS and MTG and potentially the lateral occipital complex (LOC), with the resulting signals projected back to V1–V3 via top-down feedback 38 , 39 . Thus, the observed size-contrast effect may reflect a multistage process potentially involving both early visual representations and higher-level contextual feedback mechanisms. Because the present study was designed to test retinotopic involvement rather than to disentangle early sensory contributions from feedback-related influences, we cannot determine whether the activation patterns observed in early visual cortex represent the origin of the ensemble size-contrast effect or reflect downstream effects of feedback from higher-level regions. Future work could build on these findings by employing DCM to test how early visual areas and higher-level regions interact during ensemble-level size-contrast, or by using transcranial magnetic stimulation (TMS) to test for causal contributions of different network nodes. Although the current study focuses on ensemble summary statistics, our findings point to a more general mechanism in size perception or possibly perception more broadly. Our results indicate that the visual system computes ensemble summary statistics for both attended and unattended stimulus sets and evaluates these representations relative to one another. This mechanism is consistent with judgement-based models of size-contrast illusions, in which perceived size reflects a comparison between the target and the context 12 . However, the individual strength of the ensemble size-contrast effect and the strength of the Ebbinghaus illusion measured in the screening task were not correlated, indicating that both effects do not share a single mechanism, even though both produce size-contrast-like perceptual effects. Individual-object and ensemble representations possibly share more general aspects of size coding. The consistent underestimation observed across the Ebbinghaus screening, pre-test, and fMRI session might suggest that the peripheral underestimation previously documented for individual objects 40 may also extend to ensemble summary representations. Please note that this underestimation bias cannot explain the present size-contrast effects, as these were computed as relative differences between size-contrast and size-match conditions. Beyond this, the ensemble size-contrast effect may constitute a more general mechanism by which the visual system uses ensemble statistics to code visual information efficiently. Computing and contrasting summary statistics across groups may be a basic operation underlying attentional selection. For instance, detecting a pop-out target in visual search may involve ensemble representations: the visual system may rapidly compute an average representation of the search array and identify targets as items that deviate significantly from this statistical norm, making them particularly salient 41 . The contextual modulation observed in this study, therefore, suggests that target detection in visual search may be shaped not only by individual stimulus properties but also by the statistical regularities of the surrounding context. Finally, two aspects of the present study warrant further discussion. The results revealed consistent underestimation of the average sizes of the target set of objects across all conditions. This underestimation may reflect a general perceptual bias induced by the mere presence of a distractor ensemble, independent of its average size. The present design did not include an additional target-only control condition, which would have provided a direct baseline for behavioural size judgments and neural responses to the target ensemble independent of other object sets. Based on previous studies showing accurate average-size extraction from single object sets 1 , 8 , 9 , the estimated average size in such a condition would be expected to more closely reflect the physical average size, thereby providing a direct behavioural and neural baseline for size judgments in the absence of any contextual modulation, and potentially clarifying whether the observed underestimation is attributable to distractor presence or reflects a more general perceptual bias. Additionally, we found no evidence for a differential influence of luminance on the effects of target or distractor ensemble size. The mechanism by which object groups were established is orthogonal to our central claim, and any facilitation of grouping arising from the luminance difference would actually reinforce, not weaken, the observed effect. Future studies should nevertheless examine and control this factor further by equating luminance across stimulus groups. Conclusion Our findings show that the statistical features of simultaneously presented object groups are not computed independently, but are shaped by size-contrast mechanisms. Combining behavioural results with functional imaging analyses, we found that the size of the distractor ensemble modulated the perceived average size of the target ensemble. Importantly, a similar pattern was observed for regions representing objects that participants were explicitly instructed to ignore. This mutual size-contrast effect suggests that each ensemble serves as a reference for the other, irrespective of task relevance. Together, these findings contribute to a broader understanding of how the visual system organises complex scenes and highlight the importance of contrast effects and relative coding in shaping ensemble representations. Data availability The datasets generated and analysed during the current study will be made available upon publication of the article. In the meantime, data are available from the corresponding author upon reasonable request. References Ariely, D. Seeing sets: Representation by statistical properties. Psychol. Sci. 12 (2), 157–162 (2001). Article PubMed Google Scholar Tiurina, N. A. & Utochkin, I. S. Ensemble perception in depth: Correct size-distance rescaling of multiple objects before averaging. J. Exper. Psychol. Gen. 148 (4), 728 (2019). Article Google Scholar Parkes, L., Lund, J., Angelucci, A., Solomon, J. A. & Morgan, M. Compulsory averaging of crowded orientation signals in human vision. Nat. Neurosci. 4 (7), 739–744 (2001). Article PubMed Google Scholar Haberman, J. & Whitney, D. Rapid extraction of mean emotion and gender from sets of faces. Curr. Biol. 17 (17), R751–R753 (2007). Article PubMed PubMed Central Google Scholar Alvarez, G. A. & Oliva, A. The representation of simple ensemble visual features outside the focus of attention. Psychol. Sci. 19 (4), 392–398 (2008). Article PubMed PubMed Central Google Scholar Allik, J., Toom, M., Raidvee, A., Averin, K. & Kreegipuu, K. Obligatory averaging in mean size perception. Vision. Res. 101 , 34–40 (2014). Article PubMed Google Scholar Choo, H. & Franconeri, S. L. Objects with reduced visibility still contribute to size averaging. Atten. Percept. Psychophys. 72 , 86–99 (2010). Article PubMed Google Scholar Chong, S. C. & Treisman, A. Statistical processing: Computing the average size in perceptual groups. Vis. Res. 45 (7), 891–900 (2005). Article PubMed Google Scholar Chong, S. C. & Treisman, A. Representation of statistical properties. Vis. Res. 43 (4), 393–404 (2003). Article PubMed Google Scholar Epstein, M. L. & Emmanouil, T. A. Ensemble statistics can be available before individual item properties: Electroencephalography evidence using the oddball paradigm. J. Cogn. Neurosci. 33 (6), 1056–1068 (2021). Article PubMed PubMed Central Google Scholar Memis, E., Yildiz, G. Y., Fink, G. R. & Weidner, R. Hidden size: Size representations in implicitly coded objects. Cognition 256 , 106041 (2025). Article PubMed Google Scholar Massaro, D. W. & Anderson, N. H. Judgmental model of the Ebbinghaus illusion. J. Exp. Psychol. 89 (1), 147 (1971). Article PubMed Google Scholar Murray, S. O., Boyaci, H. & Kersten, D. The representation of perceived angular size in human primary visual cortex. Nat. Neurosci. 9 (3), 429–434 (2006). Article PubMed Google Scholar Velhagen, K. & Broschmann, D. Tafeln zur Prüfung des Farbensinnes (Thieme, 1997). Erdfelder, E., Faul, F. & Buchner, A. GPOWER: A general power analysis program. Behav. Res. Methods Instr. Comput. . 28 (1), 1–11 (1996). Article Google Scholar Peirce, J. et al. PsychoPy2: Experiments in behavior made easy. Behav. Res. Methods . 51 (1), 195–203 (2019). Article PubMed PubMed Central Google Scholar RStudio Team. RStudio: Integrated development for R (RStudio, Inc, 2015). Power, J. D., Barnes, K. A., Snyder, A. Z., Schlaggar, B. L. & Petersen, S. E. Spurious but systematic correlations in functional connectivity MRI networks arise from subject motion. Neuroimage 59 (3), 2142–2154 (2012). Article PubMed Google Scholar Ciric, R. et al. Benchmarking of participant-level confound regression strategies for the control of motion artifact in studies of functional connectivity. Neuroimage 154 , 174–187 (2017). Article PubMed PubMed Central Google Scholar Wang, L., Mruczek, R. E., Arcaro, M. J. & Kastner, S. Probabilistic maps of visual topography in human cortex. Cereb. Cortex . 25 (10), 3911–3931 (2015). Article PubMed Google Scholar O’Brien, F. & Cousineau, D. Representing error bars in within-subject designs in typical software packages. Quant. Methods Psychol. 10 (1), 56–67 (2014). Article Google Scholar Benson, N. C., Kupers, E. R., Barbot, A., Carrasco, M. & Winawer, J. Cortical magnification in human visual cortex parallels task performance around the visual field. Elife , 10 , e67685. (2021). Ebbinghaus, H. Grundzüge der Psychologie Vol. 1 (Verlag von Veit & Comp, 1902). Cant, J. S. & Xu, Y. The impact of density and ratio on object-ensemble representation in human anterior-medial ventral visual cortex. Cereb. Cortex . 25 (11), 4226–4239 (2015). Article PubMed Google Scholar Jia, J., Wang, T., Chen, S., Ding, N. & Fang, F. Ensemble size perception: Its neural signature and the role of global interaction over individual items. Neuropsychologia 173 , 108290 (2022). Article PubMed Google Scholar Schwarzkopf, D. S., Song, C. & Rees, G. The surface area of human V1 predicts the subjective experience of object size. Nat. Neurosci. 14 (1), 28–30 (2011). Article PubMed Google Scholar Sperandio, I., Chouinard, P. A. & Goodale, M. A. Retinotopic activity in V1 reflects the perceived and not the retinal size of an afterimage. Nat. Neurosci. 15 (4), 540–542 (2012). Article PubMed Google Scholar Weidner, R. et al. The moon illusion and size–distance scaling—evidence for shared neural patterns. J. Cogn. Neurosci. 26 (8), 1871–1882 (2014). Article PubMed Google Scholar Im, H. Y. & Chong, S. C. Computation of mean size is based on perceived size. Atten. Percept. Psychophys. 71 (2), 375–384 (2009). Article PubMed Google Scholar Markov, Y. A. & Tiurina, N. A. Size-distance rescaling in the ensemble representation of range: Study with binocular and monocular cues. Acta. Psychol. 213 , 103238 (2021). Article Google Scholar Haberman, J. & Suresh, S. Ensemble size judgments account for size constancy. Atten. Percept. Psychophys. 83 (3), 925–933 (2021). Article PubMed Google Scholar Ortego, K. & Störmer, V. S. Similarity in feature space dictates the efficiency of attentional selection during ensemble processing. Psychon. Bull. Rev. , 1–10. (2024). Brady, T. F. & Alvarez, G. A. Hierarchical encoding in visual working memory: Ensemble statistics bias memory for individual items. Psychol. Sci. 22 (3), 384–392 (2011). Article PubMed Google Scholar Dumoulin, S. O. & Wandell, B. A. Population receptive field estimates in human visual cortex. Neuroimage 39 (2), 647–660 (2008). Article PubMed Google Scholar Amano, K., Wandell, B. A. & Dumoulin, S. O. Visual field maps, population receptive field sizes, and visual field coverage in the human MT+ complex. J. Neurophysiol. 102 (5), 2704–2718 (2009). Article PubMed PubMed Central Google Scholar Moutsiana, C. et al. Cortical idiosyncrasies predict the perception of object size. Nat. Commun. 7 (1), 12110 (2016). Article ADS PubMed PubMed Central Google Scholar Roberts, B., Harris, M. G. & Yates, T. A. The roles of inducer size and distance in the Ebbinghaus illusion (Titchener circles). Perception 34 (7), 847–856 (2005). Article PubMed Google Scholar Weidner, R. & Fink, G. R. The neural mechanisms underlying the Müller-Lyer illusion and its interaction with visuospatial judgments. Cereb. Cortex . 17 (4), 878–884 (2007). Article PubMed Google Scholar Zeng, H., Fink, G. R. & Weidner, R. Visual size processing in early visual cortex follows lateral occipital cortex involvement. J. Neurosci. 40 (22), 4410–4417 (2020). Article PubMed PubMed Central Google Scholar Bedell, H. E. & Johnson, C. A. The perceived size of targets in the peripheral and central visual fields. Ophthalmic Physiol. Opt. 4 (2), 123–131 (1984). Article PubMed Google Scholar Treisman, A. M. & Gelade, G. A feature-integration theory of attention. Cogn. Psychol. 12 (1), 97–136 (1980). Article PubMed Google Scholar Download references Acknowledgements Open access publication funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – 491111487. GRF gratefully acknowledges support by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) (Project-ID 431549029—SFB 1451). The funders have/had no role in the decision to publish or the preparation of the manuscript. We are grateful to our colleagues from the Institute of Neuroscience and Medicine for many valuable discussions. Funding This work was supported by Forschungszentrum Jülich GmbH. Open access publication funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – 491111487. Besides, GRF gratefully acknowledges support by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) (Project-ID 431549029—SFB 1451). The funders have/had no role in the decision to publish or the preparation of the manuscript. Author information Authors and Affiliations Cognitive Neuroscience, Institute of Neuroscience and Medicine (INM-3), Forschungszentrum Jülich GmbH, Leo-Brand-Str. 5, 52425, Jülich, Germany Elif Memis, Gereon R. Fink & Ralph Weidner Department of Neurology, University Hospital Cologne, Cologne University, Kerpener Str. 62, 50937, Köln, Germany Gereon R. Fink Authors Elif Memis Gereon R. Fink Ralph Weidner Contributions Elif Memis : Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Software, Validation, Visualization, and Writing- Original draft. Gereon R. Fink : Conceptualization, Funding acquisition, Methodology, Resources, Supervision, and Writing- Reviewing and Editing. Ralph Weidner : Conceptualization, Methodology, Project administration, Supervision, Software, Formal analysis, Data curation, Validation, and Writing- Reviewing and Editing. Corresponding author Correspondence to Elif Memis . Ethics declarations Competing interests The authors declare no competing interests. Additional information Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Rights and permissions Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/ . Reprints and permissions About this article Cite this article Memis, E., Fink, G.R. & Weidner, R. Object groups mutually bias each other’s perceived average size, as evidenced by neural and behavioural data. Sci Rep 16 , 31715 (2026). https://doi.org/10.1038/s41598-026-75491-3 Download citation Received : 25 February 2026 Accepted : 07 October 2026 Published : 11 October 2026 Version of record : 11 October 2026 DOI : https://doi.org/10.1038/s41598-026-75491-3 Keywords

Comments

Sign in to join the conversation

Sign In

No comments yet. Be the first to share your thoughts!

E
Written by

Editorial Team

Staff writer covering breaking news, features, and long-form analysis for NewsLive. Tracking the stories that matter most.

Stay in the loop

Get the best stories
delivered weekly

Join thousands of readers who get our top stories in their inbox every week. No spam, unsubscribe any time.