Experimental set-up for spatial DOP modulation
The experimental system combines a spatial degree-of-polarization (DOP) modulator with two independent platforms: a high-dimensional photonic neural network (PNN) and a multidimensional optical encryption–decryption system. Both platforms use scattering media and digital neural networks in different ways, as described in the following sections.
The spatial DOP modulator consists of a phase-only liquid-crystal-on-silicon spatial light modulator (SLM; Hamamatsu X13138, 1,280 × 1,024 pixels, 12.5 μm pixel pitch and 60 Hz frame rate) positioned between an input half-waveplate (HWP) and an output quarter-waveplate (QWP)–HWP pair. An expanded continuous-wave laser beam (λ = 532 nm, 250 mW) with diagonal (D) polarization, set by the input HWP, illuminates the SLM. The output QWP and HWP are mounted on high-speed motorized rotation stages and oriented at angles α and β, which are programmed together with the SLM. These polarization elements convert the phase delay ϕ applied by each SLM pixel into a state of polarization (SOP) determined by (ϕ, α, β), as described in Supplementary Note 1. The DOP and SOP of each macromode are controlled using four parameters: (δϕ, \(\bar{\phi }\), α, β).
For the two-SLM configuration (Supplementary Fig. 4), two identical Hamamatsu X15213-16L SLMs, each with 1,280 × 1,024 pixels, a 12.5 μm pixel pitch and a 60 Hz frame rate, are cascaded pixel-to-pixel using a 4f lens system. An HWP is inserted between the SLMs at a fixed angle of γ = 22.5°. The modulator is calibrated through polarimetry measurements performed with a rotating-waveplate polarimeter (Thorlabs PAX1000VIS; 0.25° accuracy), which measures S1, S2, S3 and ρ.
Spatial DOP and SOP modulation is implemented in two configurations. In the first configuration, the modulated beam is measured in a far-field plane at a distance z from the SLM. In the second, the beam is measured in the Fourier plane. Here, we focus on the far-field arrangement because it provides a versatile implementation without additional optical components. The Fourier-plane configuration, which uses a microlens array, is described in Supplementary Note 6. The working distance z is selected according to the micromode size l, which determines the diffraction distance over which micromodes mix during propagation. At full resolution (N = 32 × 32), z is approximately 5 cm. At this distance, the far-field macromode size is comparable to its size L on the SLM.
The spatial polarization modulator is validated using a non-full-Stokes polarization camera (method 1) and a full-Stokes imaging system (method 2). A Thorlabs Kiralux polarization camera with 2,448 × 2,048 pixels records the linear polarization degree \(\nu =\sqrt{{S}_{1}^{2}+{S}_{2}^{2}}/{S}_{0}\), the azimuth θ = arctan(S2/S1)/2 and the intensity S0(x, y). Full-Stokes and DOP imaging is performed by operating the camera in intensity mode and sequentially recording intensity projections through a QWP and polarizer at different orientations54. The accuracy of the spatial DOP and SOP modulation is quantified using \({\rm{RMSE}}=\frac{1}{N}{\sum }_{i}^{N}\sqrt{{\sum }_{k}|{S}_{k}^{{\rm{m}}}-{S}_{k}^{{\rm{p}}}{|}^{2}/3}\), where the superscripts ‘m’ and ‘p’ indicate measured and programmed values, respectively.
Programming the spatial DOP modulator
The active area of the SLM is divided into N square macromodes, each containing M square micromodes. Every micromode consists of l × l pixels, with l selected to fill the SLM active area for the desired spatial resolution N. For example, l = 12 pixels is used for M = 256 and N = 25 (Fig. 3d), giving a micromode size of 150 μm. The minimum macromode size required for accurate spatial DOP and SOP control is L = 25 pixels. This is achieved with M = 25 micromodes, each having a length of l = 5 pixels (Fig. 3h). For high-resolution modulation (Fig. 3i), several blank pixels with constant polarization are placed between macromodes to prevent diffraction-induced overlap.
The phase mask is generated by assigning every pixel in the jth micromode a constant phase ϕj within the range [0, 2π]. The phase value ϕj is randomly sampled from a Gaussian probability density function (PDF) associated with the ith macromode: \({{\mathcal{N}}}^{(i)}(\phi )=(1/\sqrt{2{\rm{\pi }}\delta {\phi }_{i}^{2}})\exp [-{(\phi -{\bar{\phi }}_{i})}^{2}/2\delta {\phi }_{i}^{2}],\) where the standard deviation δϕi lies between 0 and π/2 and the mean \({\bar{\phi }}_{i}\) lies between 0 and 2π. Changing \(\bar{\phi }\) moves the SOP along a trajectory on the Poincaré sphere, while the trajectory is controlled by the waveplate angles.
The modulator is calibrated by repeating the analysis shown in Fig. 2 for different combinations of (δϕ, \(\bar{\phi }\), α, β). Each data point in Fig. 2 represents a single-mask measurement. As ρ approaches zero, the polarized component becomes less defined, resulting in increased measurement uncertainty in the Stokes parameters. Averaging several statistically equivalent masks reduces the noise observed in individual measurements (Supplementary Fig. 2). The DOP is calibrated from the average modulation using the fitting function ρ = aexp(−bδϕ2) + c. The measured SOP (Fig. 2b) agrees closely with the polarization matrix model presented in Supplementary Note 2.
A mapping is then established between the Stokes parameters (S1, S2, S3, ρ) and the four control parameters (δϕ, \(\bar{\phi }\), α, β) = X. A target beam with spatially varying DOP and SOP is generated by assigning the corresponding vectors X(i) to each macromode. The effect of the number of micromodes M is investigated in Supplementary Fig. 3. In the two-SLM configuration, the waveplate angles α and β are replaced by a second tunable phase \({\phi }_{2}^{(i)}\), which is independently set for each macromode and remains constant within it. The resulting control vectors are \({X}^{(i)}=(\delta \phi ,\bar{\phi },{\phi }_{2})\). Calibration results for the two-SLM modulator are provided in Supplementary Fig. 5. All modulation patterns are generated using custom MATLAB software.
Reducing macromode crosstalk
Accurate spatial DOP and SOP modulation requires minimal interaction between neighbouring macromodes. Crosstalk can reduce modulation precision because the polarization state programmed for one macromode may influence adjacent regions. This effect is controlled by selecting an appropriate far-field distance z. The interaction between neighbouring micromodes and macromodes occurs at distances comparable to their respective diffraction lengths, zl = πl2/λ and zL = πL2/λ. Therefore, the working distance should satisfy zl ≪ z ≪ zL. This condition is readily achieved when the micromodes are much smaller than the macromodes (l ≪ L).
For example, the parameters L = 2.4 mm and l = 150 μm used in Fig. 3d–f produce an operating range of approximately 0.1 m < z < 10 m, within which macromode crosstalk is negligible. Crosstalk becomes more significant when L and l have similar values, as occurs when M is reduced to increase the number of addressable macromodes. In this regime, adjacent macromodes are separated by blank pixels. The buffer length is selected so that micromodes at the edges of neighbouring macromodes do not overlap during propagation. Residual crosstalk is measured experimentally in Supplementary Note 9. In Fig. 3i, where l = 60 μm and z = 5 cm, six blank pixels are used (d = 6). In the Fourier-plane configuration, macromode crosstalk is inherently suppressed because each microlens operates on a single macromode. This arrangement is particularly suitable for applications requiring focused, DOP-modulated illumination.
Encoding RGB colours using polarization
Encoding true RGB colour information in polarization requires modulation of both the SOP and the DOP. SOP modulation alone is insufficient because assigning colours to different points on the Poincaré sphere does not preserve the fundamental property that colours can be represented as linear combinations of primary colours. We overcome this limitation using the mapping shown in Fig. 4a: \({S}_{1}=(2{\rm{R}}-1)/\sqrt{3}\), \({S}_{2}=(2{\rm{G}}-1)/\sqrt{3}\) and \({S}_{3}=(2{\rm{B}}-1)/\sqrt{3}\), where R, G, B ∈ [0, 1]. Other colour-to-polarization mappings are also possible, including nonlinear transformations and representations based on the spherical coordinates [θ, χ, ρ] of the Poincaré sphere. The proposed approach supports RGB colour encoding with up to 8 bits per channel, limited by the SLM bit depth.
High-dimensional photonic neural network
The optical component of the high-dimensional PNN contains n random optical layers formed by a stack of n diffusers. The diffusers are Thorlabs N-BK7 ground-glass diffusers with 120-, 200-, 600- or 1,500-grit surfaces. An optoelectronic layer is implemented using a CMOS camera (Basler a2A1920-160umPRO, 1,920 × 1,200 pixels, 12-bit pixel depth) positioned 10 cm from the diffuser stack. A 4 × 4-pixel binning operation is performed directly on the CMOS sensor, providing hardware-based average pooling. The resulting 300 × 300 array of intensity values, each with 12-bit precision, is processed by a digital backend. This digital network contains two fully connected layers, 40 hidden nodes and ten output nodes corresponding to the classification categories. The final class is selected using a softmax operation, and the network is trained with the Adam optimizer.
We evaluate the high-dimensional PNN by classifying colour images from the CIFAR-10 dataset53. This benchmark contains 60,000 32 × 32 RGB images distributed across ten object classes, with 6,000 images per class. The dataset includes 50,000 training images and 10,000 test images. For comparison, we also classify corresponding greyscale images generated from the original RGB data. Each image is encoded into N = 32 × 32 polarization macromodes with a macromode size of L = 25 pixels (Fig. 4c), using the RGB-to-Stokes mapping shown in Fig. 4a.
This colour-to-polarization conversion introduces a nonlinear transformation of the input data. Phase encoding is also nonlinear55; consequently, input nonlinearity plays only a minor role in the measured performance improvement. Classification accuracy is calculated by averaging multiple independent training and testing runs. The high-dimensional PNN is modelled as a sequence of cascaded variable transmission matrices (VTMs) and partially coherent propagation, as detailed in Supplementary Note 7.
Scalability of the spatial polarization encoder
As a high-dimensional optical encoder, the spatial DOP modulator provides a resolution that scales linearly with the number of SLM pixels: N = ξ−1npx. Here, ξ = L2 = M × l2 is a set-up-dependent scaling factor. A comparison with other spatial optical encoders used in PNNs is provided in Supplementary Table 1. In our 32 × 32-macromode implementation, ξ ≈ 6 × 102 (Fig. 3h). At this scaling factor, an ultrahigh-definition 4K SLM with npx = 4,160 × 2,464 pixels could generate more than 10,000 macromodes. These results demonstrate that large-scale spatial polarization encoding can be implemented using commercially available optical components.
Multidimensional optical encryption and decryption system
Speckle-based optical encryption is used to protect polarization-encoded RGB images. A 120-grit ground-glass diffuser is placed in the focal plane of a lens with a 250 mm focal length. The transmitted speckle pattern, which serves as the ciphertext, is recorded by a CMOS camera. The measured speckle intensity is related to the input SOP through the diffuser’s transmission tensor57. The resulting 1,200 × 1,200-pixel intensity images, recorded with 4,096 intensity levels, are downsampled into vectors containing 90,000 elements.
The decryption deep neural network (DNN) contains two fully connected layers with w × N hidden nodes and 3 × N output nodes. Batch normalization and ReLU activation are applied between the layers. The hyperparameter w controls the hidden-layer width and is optimized for decryption accuracy; w = 4 is used for the results in Fig. 5. The 3 × N output vector estimates the Stokes parameters S1, S2 and S3 for the N macromodes. The decrypted image is reconstructed by applying the inverse colour mapping [R, G, B] ↔ [S1, S2, S3]. The reconstruction error produced by an incorrect mapping, which functions as a security key, is presented in Supplementary Fig. 17.
The decryption DNN used in Fig. 5 contains approximately 3.8 × 108 learnable parameters, equivalent to 12.2 Gbit at 32-bit precision. The network is trained using 20,000 plaintext–ciphertext image pairs. Decryption quality is evaluated using the Pearson correlation coefficient (PCC), an interpretable measure of image fidelity. The system can encrypt any RGB image with a resolution of up to 32 × 32 pixels. CIFAR-10 images are encrypted in Fig. 5 to demonstrate operation at the maximum supported resolution.
Source: www.nature.com


