solr icon indicating copy to clipboard operation
solr copied to clipboard

SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields

Open thiru-mg opened this issue 2 years ago • 4 comments

https://issues.apache.org/jira/browse/SOLR-16810

Description

While persisting the ManagedIndexSchema as XML, non-printable characters in field names get escaped as #nn;, where nn is the decimal representation of the non-printable character. For example, if the field name has the byte 0x14, it gets escaped as #20;. This in indistinguishable from the literal #20; in the field name. If we have two fields - one with the non-printable character and the other with the literal string, two fields get generated with the same name. Loading the resulting XML, naturally, causes an exception. To fix this, any occurrence of literal # in the field name should be escaped, with say ##. A second problem is that while escaping happens when generating XML, the corresponding unescaping does not happen on loading it.

Solution

Both the suggested fixes are here. There are tests to expose the bugs and the fixes that pass the tests.

Tests

Added two sets of tests - one for IndexSchema.java and the other for XML.java

Checklist

Please review the following and check all that apply:

  • [x] I have reviewed the guidelines for How to Contribute and my code conforms to the standards described there to the best of my ability.
  • [x] I have created a Jira issue and added the issue ID to my pull request title.
  • [x] I have given Solr maintainers access to contribute to my PR branch. (optional but recommended)
  • [x] I have developed this patch against the main branch.
  • [x] I have run ./gradlew check.
  • [x] I have added tests for my changes.
  • [ ] I have added documentation for the Reference Guide

thiru-mg avatar May 21 '23 13:05 thiru-mg

This PR had no visible activity in the past 60 days, labeling it as stale. Any new activity will remove the stale label. To attract more reviewers, please tag someone or notify the [email protected] mailing list. Thank you for your contribution!

github-actions[bot] avatar Feb 15 '24 12:02 github-actions[bot]

Thank you StaleBot..... I just checked the JIRA and I was last to chime in, so I'll take this and try and get it over the finish line.

epugh avatar Feb 15 '24 13:02 epugh

This PR had no visible activity in the past 60 days, labeling it as stale. Any new activity will remove the stale label. To attract more reviewers, please tag someone or notify the [email protected] mailing list. Thank you for your contribution!

github-actions[bot] avatar May 01 '24 00:05 github-actions[bot]