Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
Descripción general
Resumen del artículo
This paper introduces Pico-Banana-400K, a large-scale dataset of approximately 400,000 text-guided image edits, which is primarily generated and quality-controlled by AI models rather than humans. The dataset leverages Nano-Banana for diverse edit generation from real images and Gemini-2.5-Pro for automated quality assessment, providing examples for single-turn, multi-turn, and preference-based editing scenarios. It aims to establish a robust foundation for training and benchmarking the next generation of text-guided image editing models, despite inherent biases from its AI-on-AI generation and judging process.
Explícamelo como si tuviera cinco años
Computers made a huge collection of edited pictures, and then other computers decided if they looked good. This helps teach AI how to change images using simple written commands, like magic words for photos.
Posibles conflictos de intereses
All authors are affiliated with Apple. The paper explicitly states that Nano-Banana, Gemini-2.5-Flash, and Gemini-2.5-Pro models were used for dataset generation and quality assessment, which are either Apple's internal models or models developed by companies with close ties to the authors' institution. This creates a direct conflict of interest, as the authors are using and validating their employer's (or closely related entities') proprietary tools and models in the creation of a public dataset, which could have a vested interest in the dataset's perceived quality and utility.
Limitaciones identificadas
Explicación de la calificación
The paper presents a valuable large-scale dataset for text-guided image editing with a comprehensive taxonomy and robust automated quality control. However, the complete reliance on AI for both generation and judging introduces potential biases and limitations. The primary authors being from Apple, utilizing Apple's internal and proprietary models, constitutes a significant conflict of interest, warranting a reduction in the rating despite the technical contribution.
Conviene saber
Este es el análisis de Starter. Paperzilla Pro verifica cada cita, investiga los antecedentes de los autores y las fuentes de financiación, y utiliza razonamiento avanzado con IA para ofrecer información más exhaustiva.
Explorar Pro →