Matcher[source]
EDS-NLP simplifies the matching process by exposing a eds.matcher component that can match on terms or regular expressions.
Examples
Let us redefine the pipeline :
import edsnlp, edsnlp.pipes as eds
nlp = edsnlp.blank("eds")
terms = dict(
covid=["coronavirus", "covid19"], # (1)
patient="patient", # (2)
)
regex = dict(
covid=r"coronavirus|covid[-\s]?19|sars[-\s]cov[-\s]2", # (3)
)
nlp.add_pipe(
eds.matcher(
terms=terms,
regex=regex,
attr="LOWER",
term_matcher="exact",
term_matcher_config={},
),
)
- Every key in the
termsdictionary is mapped to a concept. - The
eds.matcherpipeline expects a list of expressions, or a single expression. - We can also define regular expression patterns.
This snippet is complete, and should run as is.
Patterns, be they terms or regex, are defined as dictionaries where keys become the label of the extracted entities. Dictionary values are either a single expression or a list of expressions that match the concept.
By default, matches are added to doc.ents, which is also the default input of qualifier components such as eds.negation. When matches are written only to a custom span group, pass the same group as the qualifier's span_getter:
nlp.add_pipe(eds.sentences())
nlp.add_pipe(eds.matcher(terms={"diabetes": ["diabète"]}, span_setter="mygroup"))
nlp.add_pipe(eds.negation(span_getter="mygroup"))
doc = nlp("Pas de diabète.")
assert doc.spans["mygroup"][0]._.negation is True
Alternatively, use span_setter=["ents", "mygroup"] on the matcher to expose the same matches through both locations and keep the qualifier's default input.
Parameters
| PARAMETER | DESCRIPTION |
|---|---|
nlp | The pipeline object. TYPE: |
name | The name of the component. TYPE: |
terms | A dictionary of terms. TYPE: |
regex | A dictionary of regular expressions. TYPE: |
attr | The default attribute to use for matching. Can be overridden using the TYPE: |
ignore_excluded | Whether to skip excluded tokens (requires an upstream pipeline to mark excluded tokens). TYPE: |
ignore_space_tokens | Whether to skip space tokens during matching. You won't be able to match on newlines if this is enabled and the "spaces"/"newline" option of TYPE: |
term_matcher | The matcher to use for matching phrases ? One of (exact, simstring) TYPE: |
term_matcher_config | Parameters of the matcher class TYPE: |
span_setter | Where to store matches, by default in TYPE: |
span_from_group | Whether regex spans should use the first matching capturing group instead of the full regex match. TYPE: |
Authors and citation
The eds.matcher pipeline was developed by AP-HP's Data Science team.