A synthesis of 33 papers, seven open research questions, and a practical guide to doing work that gets published.
See the open questionsThis page is a working map of the cleft lip and palate literature — what is settled, what is genuinely unresolved, and where a new researcher can contribute something real rather than another underpowered case series.
It comes from reading 33 papers spanning 2012–2026: full text for the surgical and presurgical-orthopedics cluster, abstracts plus discussion and limitations sections for the rest. Every number quoted below is traceable to a specific paper, listed in the reference table at the end.
The most important thing to notice about this literature is not any single result. It is that the strongest papers in the field are mostly negative or inconclusive, and they fail in the same way.
Four serious, well-conducted reviews, all arriving at the same place: the primary literature is too heterogeneous, too small, and too inconsistently reported to answer the questions the field actually cares about.
A field whose reviews all conclude "the evidence is too heterogeneous" is a field with structural problems that are solvable by careful work — standardizing a definition, testing a confounder, re-analyzing existing data with better methods. Those projects do not require a laboratory, a patient panel, or a grant. They require reading closely and thinking clearly, which is exactly what you can do right now.
The 33 papers fall into nine clusters. The mass is heavily concentrated in presurgical orthopedics and surgical protocol comparison.
| Cluster | Papers |
|---|---|
| Presurgical infant orthopedics | Papadopoulos 2012, Hosseini 2016, Górska 2022, Di Blasio 2024, Rabal-Soláns 2024, Bühling 2025, Machado 2026 |
| Surgical protocol & growth | van Roey 2025, Kleijnen 2023, Wu 2023, Demiröz 2021, Ameer 2023 |
| Speech & velopharyngeal function | Alighieri 2025, Tahmasebifard 2025, Liu 2026, Du 2022, Tosun 2024 |
| Feeding & growth | Kinter 2026, Ormanidou 2026, Williams 2025, Kotlarek 2026 |
| Dental anomalies | Pradhan 2020, Batwa 2018, Hamid 2025 |
| Patient-reported outcomes | Michael 2022, Kabuyaya 2024, Ruiz-Guillén 2021 |
| Epidemiology, etiology, culture | Duan 2023, Eniyew 2026, Hasanuddin 2025 |
| Perioperative care | Suleiman 2024, Grabar 2023 |
| AI & digital methods | Öztürk 2026 |
Each of these comes from a specific gap that the authors themselves named and then did not fill. That matters: a question an established author has publicly flagged as unanswered is a question you can pursue without having to argue that it is worth asking.
The tension. Four meta-analyses across fourteen years (Papadopoulos 2012, Hosseini 2016, Di Blasio 2024, Rabal-Soláns 2024) find no significant effect of presurgical orthopedics on arch dimensions, feeding, or growth. But Bühling (2025), the only prospective interventional study in this set, finds active plates outperform passive ones — 5.85 mm versus 3.57 mm of cleft width reduction — and names its own key limitation:
"the lack of numerical or categorical guidelines for the allocation of patients into the active or passive plate groups."
Bühling also reports that baseline cleft size did not predict how much change occurred. That is either a real and surprising finding or an artifact of a 20-patient sample, and it matters enormously which.
The study. Derive and externally validate a morphology-based rule for allocating infants to active, passive, or no plate. Use archived neonatal casts — every cleft center has decades of them in cupboards. Test the treatment-by-morphology interaction rather than the main effect.
Why it is feasible. Framed as heterogeneous treatment effect estimation, this sidesteps the ethical objection Machado raises about withholding treatment from newborns: you are comparing two accepted treatments, not treatment against nothing.
The tension. van Roey's most interesting paragraph is buried in his discussion. Dental arch relationships were worse under one-stage palatoplasty — but 67% of those patients came from Asia and South America, where cleft care is often not covered by insurance, against 60% of Oslo-protocol patients from Europe and Oceania. He then cites Southall: orthodontic treatment alone shifts the GOSLON score by 1.1 points, which is larger than the between-protocol differences the field has been arguing about for decades.
He raises the hypothesis. He does not test it.
The study. Meta-regression across his own 156 study groups, with country income, insurance coverage, and reported orthodontic exposure as moderators of arch relationship outcomes.
Why it is feasible. The data are already extracted and published in his supplementary files. This is weeks of work, not years, and no patients are involved. The conclusion would be outsized: it would tell the field how much of its central controversy is an artifact of who can afford a retainer.
The tension. Liu compared pharyngeal anatomy in isolated cleft palate against Pierre Robin sequence and concluded there was no difference in need ratio. But the cleft palate group had a median age of 23.7 months and the Pierre Robin group 3.8 months. The paper's own discussion concedes the age gap "likely explains" the differences found. As a comparison it is uninterpretable.
Meanwhile Tahmasebifard (2025) has the right method — MRI, 105 individuals with velopharyngeal insufficiency against 45 controls, closure height 3.7 mm lower in the VPI group — but is cross-sectional, underpowered for subgroups, and explicitly flags that "the timing and type of palatal surgery might also influence the height of VP closure. Future studies should evaluate the impact of these factors."
The study. A prospective, age-standardized preoperative imaging cohort. Measure need ratio, closure height, velar length, and cleft width at a fixed preoperative timepoint; follow to speech assessment at five years. Output: a calibrated risk score.
Why it matters. It would let surgeons match the palatoplasty technique to the anatomy rather than to habit, and it converts the Pierre-Robin-versus-isolated-cleft debate into a covariate.
The tension. Four papers touch this chain and none connect it.
The study. Does anthropometric status at the time of palatoplasty predict fistula and velopharyngeal insufficiency, independent of cleft width and surgeon experience? And the harder question: is delaying surgery to improve growth actually protective, or does it simply trade a healing benefit for a speech cost? Wu's dataset is nearly the right shape; it needs weight and length added.
The spin-off. The ethnicity finding deserves its own paper. Williams says only that "future research should investigate social determinants of health that may predict timing of palate repair." Nobody has.
The tension. This is the most methodologically distinctive question in the set, and it has teeth because ICHOM has adopted CLEFT-Q as a global standard.
The study. Formal measurement invariance and differential item functioning testing of the CLEFT-Q psychosocial subscales, with Hasanuddin's cultural-belief taxonomy as the grouping variable.
Why it matters. If the instrument is not invariant, every cross-country CLEFT-Q comparison — including the normative values everyone benchmarks against — rests on shaky ground. That is an important result whichever way it comes out.
The tension. Du (2022) ends by asking for exactly this:
"more studies are required where data mining and machine learning techniques can be implemented to quantitatively determine the characteristics of different subtypes of the malformation… the computational analysis of the 3D model where topological features such as area and length of the palatal bony structure can be automatically extracted."
Bühling (2025) has already shown that digital measurement on 3D scans agrees with manual measurement on plaster casts, so the method is validated.
Almost every quantitative claim in this literature compresses a complex three-dimensional deformity into one or two numbers — cleft width, cleft palate index, alveolar cleft width. That is very likely a major source of the heterogeneity that Rabal-Soláns and Machado both blame for their inconclusive results.
The study. Statistical shape modelling of neonatal cleft maxillae from digitized archival casts, then use the shape coordinates rather than the scalars as predictors in Directions 1 and 4.
Why it is feasible. This is the enabling method for two other directions, and the input data is already sitting in storage at every major center.
7a. A reporting-bias-corrected estimate of VPI and fistula rates. van Roey's funnel plot shows that studies reporting high rates of velopharyngeal insufficiency are underrepresented, and he states plainly that "in practice, the incidences of VPI are likely to be higher." He also notes that definitions of oronasal fistula are inconsistent or entirely absent across studies. So: apply selection models to his extracted data to estimate the true rates, and run a Delphi process to fix a definition. He even proposes the long-term mechanism — registries let complication rates be reported without tracing back to individual institutions, "which may reduce reluctance to report." That is an empirically testable claim about how researchers behave.
7b. An enhanced-recovery protocol with healing, not pain, as the primary outcome. Grabar (2023) found 61% of cleft surgeons never use a standardized perioperative protocol, but 68% would adopt one. Suleiman (2024) already supplies the evidence-based analgesia regimen and flags dexamethasone and pre-incisional infiltration as unresolved. Wu (2023) supplies the missing link: agitation and crying drive wound dehiscence and fistula. Nobody has run a cleft enhanced-recovery trial with fistula as the endpoint rather than a pain score — which is the outcome surgeons actually care about, and therefore the reason they would adopt the protocol.
Direction 2 — the access-as-confounder meta-regression. It is cheap, it uses already-published data, and it could reframe the field's central debate.
Direction 1 — presurgical orthopedics as an effect-modification problem. Seven papers in this collection have circled it for fourteen years without landing on it.
The literature contains a direct conflict on speech therapy that nobody has addressed.
They may be defining the comparator differently — Kleijnen's is "purely phonological," Alighieri's is combined — but that distinction is carrying a great deal of weight and has never been stated explicitly in print.
Alighieri's trial has 14 participants. Given that van Roey notes children in the Scandcleft trial who attended extensive speech therapy still had poor consonant accuracy, and that Alighieri's own earlier work documents parental "treatment fatigue and dropout," an adequately powered trial stratified by error type is overdue.
A short, publishable first project: a narrative review that reconciles the two, making the definitional difference explicit. Reconciling a contradiction between two respected sources is a real contribution, and it is achievable without collecting any data.
Most people entering cleft research write a retrospective chart review of their institution's last hundred cases, publish it in a mid-tier journal, and add nothing. The reviews above explain exactly why that fails: the field is drowning in small, heterogeneous, inconsistently reported single-center studies. Adding one more makes the problem marginally worse.
The way to stand out is to work on the structural problems instead. Here is how.
This is the single highest-leverage habit in research reading. Abstracts tell you what authors want to claim; limitations sections tell you what they actually know. Every one of the seven directions above came from a limitations or discussion paragraph where an author named a gap and walked away from it.
For each paper, keep three columns: known, uncertain, unknown. The third column is your research pipeline.
You do not currently have a patient panel, a laboratory, or a grant. That rules out some research — and it points directly at the kind that is both achievable and unusually valuable right now: secondary-data work. Meta-regression, reanalysis of published data, psychometric validation, systematic review, methodological critique.
Directions 2, 5, and 7a require nothing but published papers, a laptop, and statistical care. They are also among the most consequential ideas on this page. That is not a coincidence: because they require no clinical access, nobody with clinical access has bothered to do them.
Every strong paper in this collection uses the same small toolkit. Learn it once and you can read — and write — at the top of the field.
PRISMA reporting standard for systematic reviews
GRADE rating certainty of evidence
ROBINS-I / RoB 2 risk-of-bias assessment
R: meta, metafor meta-analysis and meta-regression
van Roey's paper is the best worked example: read its methods section closely and you will see all four applied properly. Then register your own protocol on PROSPERO before you start — it is free, it takes an afternoon, and it signals seriousness to any mentor you approach.
"Can I help with your research?" gets a polite no, because it creates work for the recipient. "I have drafted a protocol to test whether the association between surgical protocol and arch relationship is confounded by orthodontic access, using the 156 study groups in your 2025 meta-analysis — would you be willing to look at it?" is a different message entirely. It offers labour rather than asking for it, and it demonstrates that you read the paper properly.
The authors in this collection who are most worth writing to are the ones who publicly named gaps they had no plans to fill: van Roey, Bühling, Du, Kinter, Williams. Reference their exact sentence.
Cleft care is unusually well organized internationally, and the organizations are where the data and the collaborators are.
Being a named, familiar participant in these communities compounds over years. Start now, while you have reading time.
The temptation is to work across surgery, speech, feeding, and psychosocial outcomes at once. Resist it. The researchers who matter in this field are known for one thing — van Roey for protocol comparison, Kinter and McKinney for growth, Alighieri for speech intervention, Perry for velopharyngeal imaging. Depth is what generates invitations, collaborations, and eventually your own trial.
Given your goal of operating on children with clefts, the natural centre of gravity is the surgical cluster: Directions 1, 2, 3, and 7b all sit there, and all of them connect to decisions you will personally be making in an operating room in a few years.
Keep a running document of every gap you find, with the exact quotation and citation. Within six months of disciplined reading you will have something most junior researchers never assemble: a defensible, evidence-based map of what the field does not know. That document is the raw material for a review article, a research statement, an interview answer, and eventually a grant.
That path produces a first-author publication, a conference presentation, and two senior collaborators — without requiring a lab, a grant, or a single patient.
Every claim on this page traces to one of these papers.
| Paper | Figures cited above |
|---|---|
| van Roey 2025 | 162 articles / 156 study groups; 4,040 one-stage + 1,632 Oslo + 791 delayed-closure patients. VPI: Oslo 24%, one-stage 14%, delayed 9%. Fistula: Oslo 7%, one-stage 10%, delayed 20%. Arch-relationship sample 32.5% Asia and 34.3% South America (one-stage) against 59.8% Europe/Oceania (Oslo). Southall: 1.1-point GOSLON gain from orthodontics. |
| Kleijnen 2023 | 29 clinical questions; systematic-review evidence found for 8; 21 unanswered. |
| Rabal-Soláns 2024 | 5 studies. Alveolar cleft width mean difference −3.06 (95% CI −8.03 to 2.70), I² = 99%. Posterior cleft width −0.88 (−2.06 to 0.30), I² = 89%. |
| Machado 2026 | 207 studies mapped; nasoalveolar molding most frequent; predominance of case reports and case series. |
| Bühling 2025 | 20 unilateral cleft lip and palate + 11 isolated cleft palate controls. Mean cleft width reduction 5.05 mm; active 5.85 vs passive 3.57 mm (p = 0.024). Baseline cleft size did not affect change. Digital and manual measurement agreed. |
| Wu 2023 | 980 palatoplasties. 84.2% normal healing, 11.4% delayed, 4.4% fistula. Risk factors: surgeon seniority, widest cleft width, cleft palate index, operation time. |
| Kinter 2026 | 143 eligible studies, 46 with comparable estimates. Mean weight-for-age z-score near −1 SD; 10–50% failure to thrive or underweight. Sparse data past 12 months. |
| Williams 2025 | 414 infants (268 cleft palate, 146 with Pierre Robin sequence) across 17 teams. Repair at 13.55 vs 12.05 months. Delay predictors: Pierre Robin sequence, Hispanic ethnicity, poor growth. |
| Tahmasebifard 2025 | MRI; 105 with VPI vs 45 controls. Closure height 3.7 mm lower in VPI. No sex effect; underpowered for subgroups. |
| Liu 2026 | 29 cleft palate vs 18 Pierre Robin sequence. Need ratio 0.728 vs 0.748 (not significant), both above the normal 0.6–0.7. Median age 23.7 vs 3.8 months — the confound. |
| Du 2022 | 34 submucous cleft palate with cleft lip. No correlation between cleft lip severity and palatal bony defect. 94% male. Calls for machine learning on 3D models. |
| Alighieri 2025 | 14 children (7 per arm). One-week intensive, two sessions daily. Combined phonetic–phonological outperformed motor–phonetic; maintained at 3 months only in the combined arm. |
| Grabar 2023 | 31 of 102 surgeons responded (30.4%). 61.3% never use a standardized protocol; 67.7% willing to adopt one. |
| Suleiman 2024 | 19 randomized trials + 4 systematic reviews. Recommends suprazygomatic maxillary nerve block plus dexmedetomidine and acetaminophen/NSAIDs. Dexamethasone and pre-incisional infiltration unresolved. |
| Michael 2022 | Nigeria; adult with Veau II cleft. Post-operative scores exceeded normative values for speech and psychological function; school and social function remained below the 95% CI. |
| Kabuyaya 2024 | DRC; 43 patients aged 8–29; French CLEFT-Q before and 3 months after surgery. |
| Hasanuddin 2025 | 14 articles, 13 countries, 4 continents. Natural versus supernatural attribution of cleft causation. |
| Ormanidou 2026 | 5,651 articles screened, 296 reviewed in full. Birthweight and low birthweight by cleft type, comorbidity, and country income. |
| Kotlarek 2026 | 56.6% of academic programs include no cleft-feeding content; 81.6% of providers trained on the job. |
| Öztürk 2026 | 30 mothers; single-group pre/post with immediate post-test. Largest knowledge gain in complication monitoring, which had the lowest baseline. |
| Pradhan 2020 | 208 panoramic radiographs. 90.4% had at least one anomaly; agenesis 77.9%; lateral incisor most affected (65.2%). |
| Duan 2023 | 1,051 Tibetan patients. Cleft lip 36.44%, cleft palate 13.32%, both 50.24%. Sex ratios differ by phenotype; syndromic 3.43%. |
| Eniyew 2026 | 544 children; Bayesian bivariate multinomial model. Maternal nutrition, folate, smoking, rural residence. |
| Batwa 2018 | 96 unilateral complete cleft records. Missing teeth affected dental but not skeletal characteristics; overjet reduced with multiple missing teeth. |
I am very interested in research collaborations and initiatives to help children born with cleft lip and palate.