I have the spaces to distinguish one frame from the next. The caption must be associated with the correct frame. I prefer a little gappy over too close. I have to estimate the gap before I have it all together, as the gaps are in the image itselft. I don't have the means to adjust once it's together.
I was surprised. I was able to get the basic shot from AI rather quickly. It was one of the few images that got my very detailed instructions almost entirely correct. The only issue was the "West Coast Pride" gang sign. Just that little bit had me dancing around for hours, working some of it by hand and having 4 different AI platforms tweak it to finally get what you see. It was insane.