Authors: James T. Todd, J. Farley Norman
Categories: Article, materials, fluids, shape
Source: Journal of Vision
Doi: 10.1167/jov.25.10.11
Authors: James T. Todd, J. Farley Norman
The physical interactions among objects in the natural environment can cause dramatic changes in their shapes or patterns of motion, and those changes can provide reliable information to distinguish different types of events or materials. The present research was designed to investigate the identification of fluid materials. Observers viewed computer animations and static images of a shiny orange translucent fluid flowing from a tube into a glass jar, and they were asked to make confidence ratings about whether the depicted material looked like water/juice, oil/paint, honey/molasses, or caulk/toothpaste. The results reveal that observers can identify different types of fluid materials within broad overlapping categories based on qualitative characteristics of fluid flow that only occur within limited ranges of viscosity.
The physical interactions among objects in the natural environment can cause dramatic changes in their shapes or patterns of motion. These changes are governed in part by the intrinsic mechanical properties of objects like mass, rigidity, or elasticity. Although mechanical properties have no obvious correlates in patterns of visual stimulation, observers are able to judge them with a high degree of reliability. For example, when two visible objects collide with one another, observers can determine which one has a greater mass based on their relative velocities before and after the collision (Gilden & Proffitt, 1989; Runeson, 1977; Sanborn, Mansinghka, & Griffiths, 2013; Todd & Warren, 1982). They can discriminate the relative elasticity of objects from the trajectories of bouncing balls (Warren, Kim, & Husney, 1987), or the oscillations of a bending rod (Norman, Wiesemann, Norman, Taylor, & Craft, 2007). They can also determine the stiffness of a material from how it responds to an external force (Paulun, Schmidt, van Assen, & Fleming, 2017; Schmidt, Paulun, van Assen, & Fleming, 2017; Ujitoko & Kawabe, 2022).
Other research on the perception of mechanical properties has focused on the viscosity of fluids. Viscosity is the internal friction of a material that causes it to resist flow from external forces such as gravity. Human observers are remarkably accurate at judging the relative viscosity of fluid materials (Kawabe, Maruya, Fleming, & Nishida, 2015; Paulun, Kawabe, Nishida, & Fleming, 2015; van Assen, Barla, & Fleming, 2018; van Assen & Fleming, 2016). These judgments are based on two primary sources of information. One is the pattern of motion exhibited by a flowing liquid (Kawabe et al., 2015), and the other includes surface shape features that arise from that motion. The importance of shape features is revealed most clearly by the ability of observers to judge the viscosity of fluids from static images (Paulun, Kawabe, Nishida, & Fleming, 2015).
A particularly interesting study by van Assen & Fleming (2016) compared judgments of relative viscosity to categorical naming judgments for different types of fluid materials. They noted that many common liquids with similar viscosities, like water, milk, orange juice, or red wine, can easily be distinguished by their optical properties such as color, gloss, or opacity. Their empirical results confirmed that shape/motion features are most important for judging relative viscosity, but that observers rely primarily on optical properties for material identification.
Although optical properties are clearly important for material identification, they may not be the only relevant source of information for these judgments. It is also possible to distinguish fluids based on their dynamic properties. The viscosity of fluids varies over a huge range of possible values, but there are some qualitative characteristics of fluid flow that only occur within limited ranges of viscosity. In the present article we will argue that these qualitative characteristics can be used to define four general categories of fluids and that human observers can identify those categories with a high degree of reliability.
Low-viscosity fluids like water or juice flow easily when poured into a glass, and they can also exhibit turbulence, which creates a bumpy or chaotic surface. Medium-viscosity fluids like oil or paint can also be poured easily, but they are much more resistant to turbulence. When they flow into a glass jar, the fluid surface remains relatively flat except for a visible depression where the stream comes in contact with the contained fluid. High-viscosity fluids like honey or molasses are difficult to pour, and we often use a spoon or other utensil to speed up the process. They also diffuse quite slowly so that the liquid does not immediately flatten out when it hits the ground. Instead, it forms an amorphous pile with no visible depression where the stream comes in contact with the contained fluid. Pastes or gels are interesting materials that share some properties with both fluids and solids. Examples include caulk, toothpaste, and peanut butter. They cannot be poured at all and can only be made to flow by ejecting them from a tube (e.g., a caulking gun). The identifying factor of these materials is that they exhibit negligible diffusion. That is to say, a cross-section of the initial stream retains its shape when it contacts another surface, and it coils around itself like a rope. Unlike high-viscosity fluids, these coils form visibly distinct layers in the resulting piles.
Examples of these four general categories are shown in Figure 1. Note that the depicted materials all have the same optical properties (i.e., they are shiny, orange and translucent) but their overall surface shapes are quite different. The research described in the present article was designed to investigate the abilities of observers to identify these four general categories of fluid materials. Observers viewed computer animations and static images of liquids flowing from a tube into a glass jar like the ones shown in Figure 1. Examples of these displays can be observed in Animation S1 of the supplemental materials. All of the simulations depicted the same shiny orange translucent material, but their viscosities varied over five orders of magnitude. The response task was similar to those used by Todd and Norman (2019); Todd and Norman (2020) and Norman, Todd, and Phillips (2020) to study the identification of metal and dielectric materials. For each display, observers made confidence ratings about whether the depicted material looked like water/juice, oil/paint, honey/molasses, or caulk/toothpaste. The results show clearly that observers can identify these categories with a high degree of reliability but that they do not have sharply defined boundaries.

The stimuli were displayed on an Apple iMac computer with a 24-inch 4.5K retina display (4480 × 2520 pixel resolution). The monitor was located at a 60 cm viewing distance. The luminous intensity of the monitor, measured over an area of 25°, had a minimum intensity (for black) of 1 cd/m^2^ and a maximum intensity (for white) of 136 cd/m^2^.
The displays were judged by one of the authors (JFN) and 10 other observers who were completely naïve about the purpose of the experiment or how the displays were generated. All observers possessed normal or corrected-to-normal visual acuity.
The animations were created using the Phoenix fluid simulator by Chaos. Phoenix is a grid-based simulator that computes the average velocity and density of particles within small regions of a scene called voxels, and it incrementally tracks how the velocity in each region changes over time based on the Navier-Stokes equations. The grid dimensions for these simulations were 89 × 89 × 102 cm, and the voxel size for the computations was 0.35 cm.
The displays depicted a shiny black tube that was inserted into a clear glass mason jar. The tube had a diameter of 16 cm, and its height extended above the vertical edge of the display area. The jar had a diameter of 87 cm and a height of 93 cm. It rested on a flat ground surface that gradually curved upward behind it. This is typically referred to in photography as an infinity curve or an infinity cove. The scene was illuminated by a 100×100 cm area light that was directly above the glass jar at a distance of 330 cm from the ground.
Fluid was emitted from the lower end of the tube at a velocity of 40 cm/sec. At the beginning of each animation, the tube was positioned just above the bottom of the jar, and it gradually moved upward in a helical trajectory with a diameter of 38 cm. There were 13 possible fluid viscosities of 10^0^, 10^0.16^, 10^0.33^, 10^0.5^, 10^1^, 10^1.5^, 10^2^, 10^2.5^, 10^3^, 10^3.5^, 10^4^, 10^4.5^, and 10^5^ centipoise (cP). To provide some frames of reference, water, canola oil, honey, and toothpaste have viscosities of 10^0^, 10^1.5^, 10^3.5^, and 10^5^ cP, respectively.
The animations consisted of 200 individual frames with a spatial resolution of 800 × 800 pixels. They were computed using the V-Ray renderer by Chaos, and they were presented at a rate of 30 frames/sec. The rendered images were globally tone mapped for the Apple monitor into the sRGB 2.1 color space with a D65 white-point and a gamma of 2.2. No other global histogram adjustments (e.g., tint or burn) or local sharpening or contrast enhancement operators were used. The entire set of stimuli included the 13 animations plus the four static images shown in Figure 1. These were obtained from the 200^th^ frame of the animation sequences with viscosities of 10^0^, 10^1.5^, 10^3.5^, and 10^5^ cP, respectively. All stimuli were presented for a duration of 6.67 seconds.
On each trial, observers were presented with a computer animation or a single image. Each animation had a duration of 6.67 sec and was followed by a blank screen. Observers were required to categorize the depicted material by adjusting four sliders with a hand-held mouse (see Todd & Norman, 2019; Todd & Norman, 2020; Norman et al., 2020). Each of the sliders represented a different category labeled water/juice, oil/paint, honey/molasses, or caulk/toothpaste, and a digital readout was also provided for each one. Observers were instructed to adjust the sliders to indicate their confidence rating for each of the four possible categories, and that these confidence ratings should always sum to 100%. Any set of responses that did not satisfy this criterion would prevent the program from advancing to the next trial. Because of the digital readout of their settings, observers had no difficulty conforming to this instruction.
At the beginning of each experimental session, the details of the response task were explained, and the observers were shown examples of real fluids poured into a small glass. These liquids included water, canola oil, honey, and caulk. Because caulk cannot be poured, it was ejected into the glass using a caulking gun. During each block of trials, the 17 possible stimuli (13 animations and four static images) were presented in a random order, and all observers participated in two blocks.
The average ratings for each display are shown in Figure 2. The graph in the left panel shows the results obtained for the animated stimuli. Note how the ratings change systematically over varying levels of viscosity. When the depicted material had a viscosity of 10^0^ centipoise (cP), it was rated as water/juice with almost 100% confidence, but those ratings dropped precipitously as viscosity was increased. The confidence ratings for oil/paint reached a maximum value of 76% when the simulated viscosity was 10^1.5^ cP, and they decreased systematically for viscosities above and below that critical value. The honey/molasses judgments peaked at 87% when the viscosity was 10^3^ cP, and the caulk/toothpaste ratings peaked at 77% at the highest simulated viscosity of 10^5^ cP. One interesting aspect of these results is that the observers’ responses were close to veridical. That is to say, the peak response for each of the four categories occurred within the measured range of viscosities for those materials.

The standard errors of these ratings ranged from 0% to 8.3%. The lowest variability occurred when the average confidence ratings were close to the minimum and maximum values of 0% and 100%, and the highest variability occurred at the perceived boundaries among the different categories. The average standard error over the entire set of ratings was 3.0%. The average test-retest correlation between the first and second session for each observer was 0.81, and the average correlation between each pair of observers was 0.72. These findings indicate that there were some individual differences among the different observers. For example, one observer only used ratings of 0% and 100% without any intermediate values, and another one had 0% confidence that any of the depicted materials looked like caulk or toothpaste. On the whole, however, the ratings of different observers were highly consistent with one another.
The results were quite similar for the static displays shown in the right panel of Figure 2, although the average standard error increased to 3.2%. Note that the peak confidence ratings were comparable for the moving and static displays except for the ratings of honey/molasses, which were about 30% lower in the static condition. We suspect this may be due to the fact that the two helical turns of the stream blend together into a single pile, which, in the absence of motion, could easily be mistaken for a single turn of a more viscous material. It is important to keep in mind that all of the stimuli had exactly the same optical properties of color, gloss, and opacity. Thus the fact that observers can categorize these materials from static images indicates that there is useful visual information from the shapes of the fluid surfaces (see also Paulun et al., 2015).
The displays used in the present experiment were created by modeling fluid materials as large systems of interacting particles. By controlling the forces that attract and repel these particles at a microscopic level, it is possible to simulate the macroscopic behavior of a wide range of materials. The simulator we used computes the average velocity and density of particles within small regions of a scene called voxels, and it incrementally tracks how the velocity in each region changes over time based on the Navier-Stokes equations. The initial conditions for these simulations involve a large number of parameters, including the voxel size, the time steps per frame, fluid viscosity, surface tension, stickiness, droplet size, the rate of diffusion, and the initial stream velocity from the source.
Bates, Yildirim, Tenenbaum, and Battaglia (2019) have argued that human observers perform a similar incremental simulation of particle systems to predict the ongoing behavior of fluid materials. This is consistent with a mentalistic approach to event perception that has become popular during the past decade. Proponents of this view argue that the human mind is able to predict the dynamic interactions among objects using an internal “geometry engine” that has somehow been endowed with a complete knowledge of Newtonian mechanics (Battaglia, Hamrick, & Tenenbaum, 2013; Hamrick, Battaglia, Griffiths, & Tenenbaum, 2016; Schwettmann, Tenenbaum, & Kanwisher, 2019; Ullman, Spelke, Battaglia, & Tenenbaum, 2017). However, there are several serious problems with that argument. For example, how could observers possibly know that the behavior of fluid materials can be decomposed into the rigid interactions of thousands of small particles that are not visible to the naked eye? How is it possible to learn solutions to the Navier-Stokes equations from occasional observations of fluid materials, and how could the internal geometry engine determine the many free parameters that would be necessary to set the initial conditions for performing a mental simulation? Unfortunately, the proponents of a mentalistic approach say nothing at all about how to address these issues.
A more plausible explanation of the perception of fluid dynamics has been suggested by researchers at the University of Giessen in Germany and the NTT communications science laboratory in Japan (Kawabe et al., 2015; Paulun et al., 2015; van Assen & Fleming, 2016, van Assen et al., 2018). They argue that the perception of fluid dynamics is based on a set of motion or shape features that are exhibited by fluid materials. It is important to note that a feature-based account of dynamic event perception does not require any knowledge about Newtonian mechanics, and it also makes the problem of learning much easier. The basic idea is that observers seek out potential sources of optical information that allow them to distinguish different types of events that they encounter in the natural environment.
Some sources of information about fluid dynamics vary continuously over changes in viscosity. The most obvious of these is the rate of flow. Materials with low viscosity like water flow at a much faster rate than materials with high viscosity like honey. Note that this effect may also be visible in static images. When they are photographed at equivalent moments in time, less viscous fluids flow farther than more viscous fluids (Van Assen et al., 2018). A particularly salient aspect of flow rate is the time it takes for a fluid to reach the ground when it is poured or ejected from a tubular source. We attempted to eliminate that information in the present displays by initially placing the tubular source close to the bottom of the glass jar and then gradually move it upward. However, there was still some information about flow rate in the speed at which the material covered the bottom of the glass container. For example, the materials depicted in the left panels of Figure 1 completely covered the bottom of the glass jar during the 6.67-second display duration, whereas those depicted in the right panels did not.
Let us now consider how it might be possible to measure the motion in these displays. Kawabe et al. (2015) computed the luminance flow fields for fluid animations using an iterative pyramidal Lucas–Kanade method (Bouguet, 2001; Lucas & Kanade, 1981). All algorithms for computing luminance flow (see also Horn & Schunck, 1981) were originally designed to deal with a moving camera (or observer) within a fixed visual scene where all surfaces have Lambertian reflectance and a fixed spatial relationship with the sources of illumination. Under those specific conditions, the luminance at any point on a surface remains invariant over time, and the luminance flow is identical to what would be produced by the projected motions of patterns of surface texture. However, if an object moves relative to the sources of illumination, then the luminance at each surface location will change over time. Thus movements of an observer relative to a fixed object produce very different patterns of luminance flow than movements of an object relative to a fixed observer (see Todd, 1985).
For example, the left panel of Figure 3 shows a smoothly shaded image of a curved surface, and the right panel shows the pattern of iso-luminance contours in that image that connect points with identical luminance. Patterns of iso-contours (which are also referred to as level sets) are a particularly convenient way of representing the structure of the luminance field, while stripping away the relevant information for the perception of three-dimensional shape from shading (see Norman & Todd, 2025). They can be generated quite easily in an image editing application (e.g., Photoshop) by posturizing an image into bands of uniform intensity and then applying an edge filter. Animation S2 in the supplementary materials shows several videos of how these displays change over time as the camera rotates about the depicted object, or the object rotates relative to a fixed camera. Note that when the camera moves, the iso-luminance contours appear to rotate rigidly in depth relative to the point of observation, and that provides a strong perception of three-dimensional structure from motion. However, when the object rotates relative to the fixed camera (and the sources of illumination) the iso-luminance contours are deformed in a very different manner with numerous sources and sinks that are incompatible with rigid motion in depth. The luminance flow field in that case does not provide a valid representation of how the actual object is moving relative to the observer.

Figure 4 shows the iso-luminance contours for the four images presented in Figure 1, and Animation S3 in the supplemental materials shows videos of how those contours deform over time. Note that all of these deformations have a high degree of complexity with numerous sources and sinks where contours appear and disappear. These deformations involve several distinct components that are produced by smooth occlusion contours, diffuse (Lambertian) reflectance, and specular highlights, all of which transform in a different manner based on the orientation of the fluid surface relative to the observer and the sources of illumination (see Todd, 1985). Dynamic changes in the gradients of shading caused by translucency are also influenced by the volume of the material that transmitted light must traverse to reach the point of observation. None of these patterns of deformation have been systematically analyzed, so it is difficult to determine if any of them provide useful information for the perception of fluid materials.

We suspect that the deformations of smooth occlusion contours (also referred to as the rim) may be the most informative aspect of motion in these displays. The rim is defined mathematically as the locus of points whose surface normals are all perpendicular to the line of sight. An important subset of the rim is the silhouette that is defined as the locus of points that separate a surface from its background, and it is often easy to identify by segmenting the colors in an image. For example, the lower row of Figure 5 shows the silhouettes of fluid materials in the last frame of the animation sequences for four different viscosities, and animation S5 in the supplemental materials shows how these silhouettes deform over time. Note that the gradual expansion of the silhouette provides a relatively reliable measure of the rate of fluid flow without the confounding factors that are inherent in deformations of image shading (i.e., the luminance flow field).

The fact that observers can identify fluid materials from static images indicates that motion is not a necessary source of information for perceptual identification. Fluid materials can also be distinguished by their optical properties such as color, gloss, or opacity (van Assen & Fleming, 2016), or by the shape of the fluid surface (Paulun et al., 2015). The present investigation was focused primarily on aspects of shape that change qualitatively as a function of increasing viscosity. Our results suggest that fluids can be perceptually subdivided into four general categories based on specific shape features that only occur within limited ranges of fluid viscosity.
The first of these categories is sometimes referred to as non-viscous or low-viscosity fluids, and it includes materials like water, juice, milk, and wine. Low-viscosity fluids flow easily when poured into a glass, and they also can exhibit turbulence, which creates a bumpy or chaotic surface (see left panel of Figure 1). As the turbulence decreased in the present experiment, the observers’ categorization judgments gradually shifted from water/juice to oil/paint.
Medium-viscosity fluids like oil or paint can also be poured easily, but they are much more resistant to turbulence. When they flow into a glass jar, the fluid surface remains relatively flat except for a visible depression where the stream comes in contact with the contained fluid. It is important to recognize that viscosity is not the only factor that can reduce or eliminate turbulence. That can also be achieved by increasing the surface tension of a low-viscosity fluid or by reducing the stream to a drizzle.
High-viscosity fluids like honey, molasses, or caramel are difficult to pour, and they also diffuse quite slowly so that a stream of liquid forms an amorphous pile with no visible depression where it comes in contact with the contained fluid. However, these fluids still exhibit diffusion over a sufficiently long time scale. A pile of honey will eventually diffuse into a flat surface once the flow of liquid is turned off just like the turbulent structure of a water surface. Low-, medium-, and high-viscosity liquids are impossible to distinguish once they reach a state of equilibrium. Thus they can only be identified perceptually when observed in a state of flux.
Pastes and gels, like caulk, toothpaste or peanut butter cannot be poured at all. They can only be made to flow by ejecting them from a tube (e.g., a caulking gun). The identifying factor of these materials is that they retain the shape of the nozzle from which they were ejected. This is because the forces that resist shape change are much greater than the forces that cause diffusion. If you eject a stream of caulk into a jar, it will retain its shape indefinitely. This causes the flowing material to form distinct layers when the stream piles up on itself, as is especially noticeable in the right panel of Figure 1.
It is important to keep in mind that the design of the present experiment forced the observers to limit their judgments to four possible named categories. The primary motivation for this approach is that we could identify a set of diagnostic features for each one. The fact that the observers could perform this task with high test-retest and interobserver reliability indicates that these are indeed natural distinctions for the perception of fluid dynamics. However, it is possible that there are other features we have failed to notice that could potentially allow a more fine-grained categorization.
There is one other shape feature we have not yet mentioned that can also be used to distinguish certain types of fluids. If a fluid is mixed with a gas (e.g., carbon dioxide), it can cause the formation of bubbles or foam, which is a salient visual feature. The presence or absence of foam is what allows us to visually perceive the difference between carbonated and uncarbonated water or the difference between dark beer and coffee.
Van Assen and Fleming (2016) have highlighted how the optical properties of fluids like color, gloss, and opacity provide important information for the visual recognition of fluid materials. This is what allows us to distinguish among fluids with the same viscosity such as water, milk, juice, or wine. When optical properties are combined with the patterns of shape and motion that arise from dynamic events, it provides a rich alphabet of perceptually distinct features that can be used to identify a wide range of materials.
Separating the effects of perceived shape and optical properties can be difficult because they frequently interact with one another. For example, consider the four images shown in Figure 1. All of those images were generated with the same reflectance and opacity parameters, but the one on the left with the lowest viscosity appears more translucent. This is because the turbulence produces thin sheets of liquid that light can pass through with only minimal filtering from the pigmented particles in the volume of the material (Marlow & Anderson, 2021; Marlow, Kim & Anderson, 2017). Could this have affected observers’ judgments in the present Experiment? The top row of Figure 5 shows images of an opaque material with different levels of viscosity, and Animation S4 of the supplementary materials show how those displays deform over time. Note that the diagnostic shape features for distinguishing these materials stand out even more when the effects of translucency are removed.
The bottom row of Figure 5 shows images of a black material that reflects no light at all, and Animation S5 of the supplementary materials shows how those displays deform over time. Even though all the shading gradients have been removed, there is still some residual information in the structure and deformation of the fluid silhouette. Note how the turbulence of the low viscosity material is revealed by the fractal structure of the upper boundary, and how the layering of the stream is clearly visible in the simulated paste. The high-viscosity (honey) material appears quite similar to the paste in the first helical cycle of the ejection tube, but in the second cycle the two layers merge into one. Although this merging of layers is not visible in the static image, it is clearly evident in the pattern of motion.
One way of interpreting current research on the perception of fluid dynamics is that the use of shape and/or motion features provide a heuristic for computing the physical property of viscosity (e.g., see Gilden & Proffitt, 1989; Todd & Warren, 1982). However, we think that interpretation is highly misleading. The concept of a heuristic implies that observers can only approximate actual properties of events as defined by physicists. The problem with that is that few observers have any formal training in physics, so how could they know anything at all about mathematically defined concepts in that field? An alternative way of interpreting these results is that shape/motion features are the sources of optical information that allow us to distinguish different types of events, and that human language has evolved a set of names to describe the properties of those events, and to divide them into perceptually identifiable categories. Although the perceived properties of these events may have a superficial relationship with the more formal concepts of physics, they are by no means equivalent.
Schmid and Doerschner (2018) conducted an interesting study on the perception of falling objects that break apart upon contact with the ground. Whereas hard body objects shatter into small rigid pieces, soft body objects splatter into small deformable pieces. Observers were asked to rate each event with respect to a wide range of material attributes, and several of those such as hard, soft, mushy, fluid, or gelatinous were highly correlated among the ratings of different observers. Because none of these terms has a formal definition in physics, they are best described as perceptual attributes or dimensions. These observations pose a serious problem for models of perception that rely on an internal geometry engine to simulate fluid dynamics at a microscopic scale, because those models have no way of converting the internal parameter values into distinct psychological dimensions at a macroscopic scale.
The research described in the present article was designed to investigate the identification of fluid materials. The results reveal that observers can perceptually distinguish different types of fluid materials within broad overlapping categories based on qualitative characteristics of fluid flow that only occur within limited ranges of viscosity. Low-viscosity fluids like water or juice can be identified by the presence of turbulence. Medium-viscosity fluids like oil or paint exhibit smooth laminar flow with a visible depression where the stream comes in contact with the contained fluid. High-viscosity fluids like honey or molasses produce amorphous piles when they are poured into a container; and pastes like caulk or toothpaste form distinct layers as a stream coils around itself like a rope.