Skip to content

Add independent filters, optional destinations and python_values to queries - #77

Open
srmnitc wants to merge 6 commits into
audit-fixesfrom
query-features
Open

srmnitc wants to merge 6 commits into
audit-fixesfrom
query-features

Conversation

@srmnitc

@srmnitc srmnitc commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Stacked on #76. The base is audit-fixes, so this diff shows only these 6 commits. Once #76 is merged and its branch deleted, GitHub will retarget this PR to main.

Adds the three query features whose absence meant queries against the atomRDF knowledge graph were being written as raw SPARQL.

onto.query(kg, c.AtomicScaleSample,
           [c.hasNumberOfAtoms > 1000,          # several independent filters,
            c.hasNumberOfAtoms < 5000,          # e.g. a range on one property
            c.hasSpaceGroupSymbol.optional],    # keep samples without this value
           python_values=True)                  # Python values, numeric dtypes

Several independent conditions

A query could hold only one condition (ValueError: Only one condition is allowed), so [natoms > 50, volume < 60] or a range [x > 1, x < 5] had to be squeezed into a single & expression. Now:

  • Each condition becomes its own FILTER, which SPARQL combines with AND.
  • A term used twice (e.g. a range on one property) appears once in SELECT.
  • The operands of an &/| expression that are added only for their paths no longer carry their own condition, so an OR is never narrowed to an AND.

test_prepare_destinations_multiple_conditions_error asserted the old limit, so it is replaced by a test that accepts several conditions.

Optional destinations: term.optional

Every destination was required, so a sample missing any one property dropped out of the results. term.optional has the same shape as .any / .only:

  • The destination's triple patterns, type constraint and condition go into their own OPTIONAL { } block.
  • A condition on an optional value therefore only leaves the value unbound; it does not remove the row.
  • It also works on the last step of a stepped query: [hasCell, volume.optional].

_get_triples keeps its signature. The per-destination triple patterns it now builds on are available from _get_triples_per_destination.

query(..., python_values=True)

Result cells are rdflib terms in object columns by default. With python_values=True:

  • Literals become Python values (via toPython()) and IRIs become strings.
  • Numeric columns get a numeric dtype, with NaN for missing optional values.

The default is unchanged, so atomRDF, which passes its arguments through, is not affected.

Also included

  • Fix: stepped queries whose intermediate step is an object property, e.g. [[hasMaterial, hasAltName]], raised IndexError (this also happens on main). When the path was merged, the property's range class was appended as well, shifting later nodes into the wrong triple positions.
  • Docs: a section for each feature in examples/02_advanced_query_building.ipynb, with outputs; existing cells are unchanged.

Verification

  • 196 tests pass. The new tests are in tests/test_features.py, and each fails without its change.
  • The generated queries parse with prepareQuery(..., initNs={}).
  • A live query against https://matkg.pyscal.org/sparql combining all three features returns the expected rows, with an int64 column for hasNumberOfAtoms.
  • Notebooks 01, 02 and pizza run against this branch.

A query could hold only one condition ('Only one condition is allowed'),
so [natoms > 50, volume < 60] or a range [x > 1, x < 5] failed. Each
condition now becomes its own FILTER, which SPARQL combines with AND, and
a term used twice is selected once. Operands of & / | that are appended
for their paths no longer carry their own condition, so an OR is never
narrowed to an AND.
Every destination was required, so a sample missing any one property
dropped out of the results. term.optional (like .any / .only) marks a
destination whose triple patterns, type constraint and condition are
wrapped in an OPTIONAL block; a condition on an optional value therefore
only leaves it unbound instead of removing the row.
[[hasMaterial, hasAltName]] raised IndexError: when merging the path from
an object property, its range class was appended as well, shifting every
following node into the wrong triple position. The object of the
property now takes the place of the range class.
Result cells were always rdflib terms in object columns, so every value
had to be converted before plotting or arithmetic. With
python_values=True literals become Python values and IRIs strings;
numeric columns get a numeric dtype, with NaN for missing optional
values. The default is unchanged.
Condition and type lines inside an OPTIONAL block were indented less than
its triple patterns.
Adds a section for each to the advanced query building example, with
outputs, and lists them in its summary. Existing cells are unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant