Back to timeline

Research · July 7 – 12, 1997

Lexicalised statistical parsing

At the joint ACL and EACL conference in Madrid (7-12 July 1997) Michael Collins of the University of Pennsylvania described three generative models of lexicalised context-free grammar. On section 23 of the Wall Street Journal in the Penn Treebank the best reached 88.1% constituent precision and 87.5% recall.

Why it matters

A statistical parser trained on nothing but Penn Treebank trees came close to 90% correct constituents and at the same time began recovering subcategorisation and wh-movement, which are needed to read off a tree who did what. It was measured on the same split as its predecessors, so the progress shows in the figures: Magerman's 1995 parser had 84.3 and 84.0%.

Training used sections 02-21 of the Wall Street Journal in the Penn Treebank (about 40,000 sentences), testing section 23 (2,416 sentences), with PARSEVAL measures. On all sentences up to 100 words Model 2 has 88.1% precision and 87.5% recall; Collins's previous model of 1996 85.7 and 85.3; Magerman's of 1995 84.3 and 84.0. On sentences up to 40 words (2,245) Model 2 gives 88.6 and 88.1. Model 3 recovers movement traces with 93.3% precision and 90.1% recall on 436 cases. Every constituent has a head word on which the probabilities of its parts depend. What the record does not claim. The figures 88.1 and 87.5 are easily swapped: in the paper 88.1 is precision and 87.5 recall. Charniak also built a lexicalised generative model in 1995, and Collins cites it, so the model is not the only one of its kind.

Event record

Event date
July 7, 1997 – July 12, 1997
Timeline date
Event date
Verification
Sources gathered automatically · September 26, 2026
Lines
ID
evt-0846

The 35th Annual Meeting of the ACL with the 8th Conference of the EACL, Madrid: the ACL Anthology gives the month, Crossref the days, 7-12 July 1997.

Sources

Related events