Thursday, January 10, 2019

Incomplete editorial fingerings

Say we have the following fingerings in our gold standard:
  1. 1234 131 (1 annotator vote) .1
  2. 1xx4 1x1 (3 votes) .3
  3. 12xx x3x (2 votes) .2
  4. 1x1x 1xx (4 votes) .4
How much credit should a model get for suggesting 1234 131? Or 1214 131? That is, how likely is it that the user will accept the advice? How confident are we that the advice is likely to be good?

Do we just sum over all the matches? So 1234 131 would get .1 + .3 + .2 = .6? And 1214 131 would get the same: .2 + .4 = .6? But don't we have more evidence that 1234 131 is a good fingering?

We could amplify a fingering sequences based on how many actual annotations it has:
  1. 1234 131 (7 notes x 1 annotator = 7 votes) .189
  2. 1xx4 1x1 (4 x 3 = 12 votes) .324
  3. 12xx x3x (3 x 2 = 6 votes) .162
  4. 1x1x 1xx (3 x 4 = 12 votes) .324
Or I am thinking the more specific sequences should be amplified by less specific sequences that do not contradict them:
  1. 1234 131 (1 + 3 + 2 = 6 votes) .353
  2. 1xx4 1x1 (3 + 2 = 5 votes) .313
  3. 12xx x3x (2 votes) .125
  4. 1x1x 1xx (4 votes) .25
Or a combination, like this:
  1. 1234 131 (7 + 12 + 6 = 25 votes) .4
  2. 1xx4 1x1 (12 + 6 = 18 votes) .3
  3. 12xx x3x (6 votes) .1
  4. 1x1x 1xx (12 votes) .2
Shouldn't #3 and #4 reinforce each other somehow?

Also, shouldn't the amplification run both ways? If a complete annotation agrees with a partial, what does that tell us? An editor, who is pitching his advice to a general audience in what we assume is a minimally idiosyncratic way, says these milestone fingerings are the most important.

Given this, I am leaning back toward the simple summing approach I mentioned first, but amplifying by digit. So 1234 131 would get .189 + .324 + .162 = .675, and 1214 131 would get .162 + .324 = .486. This at least passes the smell test.

It strikes me that the sparseness of editorial scores may actually be a blessing. All of my trevails with edit distances seem moot in editorial scores, or at least rendered less pertinent. The editor implicitly tells us which specific annotations are most important and which are free to vary. Using Hamming distance here is less controversial: a digit either matches or it doesn't. This is clearly justified if we assume the editor is being explicit in the advice that actually appears.

Are we really not striving to model the behavior of a good editor and less that of a good pianist?

Tuesday, January 8, 2019

The semantics of editorial fingering

Dear Anne and Justin:

Fingering annotations are typically sparse in editorial scores. This is the case even in pedagogical works intended for beginners. Editors understandably do not want to clutter their scores with unnecessary information. However, this state of affairs makes it difficult to leverage editorial scores as sources of fingering data suitable for training and evaluating computational models.

This raises a number of questions for me.

What is the intent of the editor? Is it to provide complete guidance in a compact format, as appears to be the case in beginner scores? (I remember being puzzled and a little irritated by the missing annotations. Why did I have to interpolate when I don't know what I am doing?!) Or is it to convey major transitions only (hand repositionings) and leave other "minor" decisions to the performer? Or is it to provide advice only in areas of special difficulty and to leave the rest to the performer's discretion? Or is it a combination of these intents, with the emphasis varying over the length of a piece?

How much do you think two pianists would agree when transforming a typical sparsely annotated score into a completely annotated score? Would this vary by editorial intent? (We have data we could use to tease out some answers here, I think. In the WTC corpus, we should have complete human data overlapping sparse editorial data that agree on the notes marked by the editor.)

Are there rules you apply to "fill in the blanks" in editorial scores? Are such rules discussed or codified in the pedagogical literature? Would doing so constitute a potential contribution to the pedagogical literature? Is this something you would like to pursue on its own merits?

Some time ago, I tried to brainstorm on this topic. ("MDC"--Manual Data Collector--is what I originally called the abcDE editor.)

Filling in the blanks would definitely be in order to augment our training data for machine learning. But the existence of blanks is also cramping my style in validating my latest novel evaluation metric. (It involves clustering fingering advice according to how "close" the individual fingering suggestions are to each other. This idea of closeness, already somewhat controversial, is even more strained when the suggestions are riddled with blanks.)

Thursday, July 19, 2018

2018-07-19 status

Done

Administrivia

  1. Booked trip to ISMIR 2018 in Paris.

Model Building

  1. Finailized "Corrected Parncutt" implementation for everything by cyclic patterns.
  2. Confirmed inconsistencies in published Parncutt results:
    • Small-Span and Large-Span penalties are conflated.
    • Small-Span penalty definition is inconsistent.
    • Position-Change-Count and Position-Change-Size penalties are incorrect.
    • Penalty totals are incorrect.
    • Explanatory example has confusing/incorrect costs.
  3. Completed full regression test of code base.
  4. Met with Alex Demos and agreed to co-author paper for Music Scientiae on Parncutt, Corrected Parncutt, Improved Parncutt, and how to tell them apart. Will dry run some of this material in late-breaking paper at ISMIR.

    Doing

    1. Implementing support for, and clarifying definition of, cyclic pattern constraint in Parncutt. (Should also do this for Sayegh and Hart.)
    2. Double-checking pruning mechanism in Parncutt.
    3. Writing up findings on "Corrected Parncutt" model, initially for ISMIR submission.
    4. Adding mechanism to learn weights for "Improved Parncutt" rules from training data.

    Struggling

    1. How does one compare two ranked lists of sequences to a third and claim one of the two is more similar to the third in a statistically significant way? That is the big question.

    In Scope

    1. Implementing crude automatic segmenter.
    2. Developing staccato/legato classifier.
    3. Demonstrating improved Parncutt via #1 and #2.
    4. Debugging Sayegh model, which produces results inconsistent with training data.
    5. Developing better test cases for Sayegh.
    6. Updating abcDE to support manual segmentation.
    7. Completing and polishing abcD for entire Beringer corpus.
    8. Defining initial benchmark corpora and evaluation methodology.
    9. Implementing convenience methods for reporting benchmark results.
    10. Moving Beringer corpus to MySQL database.
    11. Enhancing Parncutt, following published techniques and pushing beyond them.
    12. Enhancing Hart and Sayegh to return top n solutions.
    13. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
    14. Adding support to abcDE for annotating phrase segmentation.
    15. Debugging Dactylize 88-key circuit.
    16. Collecting fingering data from JB performances in Elizabethtown.
    17. Completing Dactylize II circuit.
    18. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
    19. Defining procedure for sanity test of production automatic data collector (including Beringer data).
    20. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
    21. Implementing end-to-end machine learning experiment, using Beringer abcD data.
    22. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

    Friday, June 22, 2018

    2018-06-22 status

    Done

    Model Building

    1. Re-implemented Parncutt cost functions for both hands.
    2. Identified possible inconsistencies in published Parncutt model description.
    3. Added feature to track more granular cost details in Dactyler models.
    4. Tracked individual rule costs to facilitate analysis of Parncutt results.
    5. Completed successful regression test for hacked-up Didactyl code. The APIs they are a-changin'.

      Doing

      1. Testing Parncutt cost functions.
      2. Reproducing original Parncutt results.
      3. Adding mechanism to learn weights for Parncutt rules from training data.

      In Scope

      1. Implementing crude automatic segmenter.
      2. Developing staccato/legato classifier.
      3. Demonstrating improved Parncutt via #1 and #2.
      4. Debugging Sayegh model, which produces results inconsistent with training data.
      5. Developing better test cases for Sayegh.
      6. Updating abcDE to support manual segmentation.
      7. Completing and polishing abcD for entire Beringer corpus.
      8. Defining initial benchmark corpora and evaluation methodology.
      9. Implementing convenience methods for reporting benchmark results.
      10. Moving Beringer corpus to MySQL database.
      11. Enhancing Parncutt, following published techniques and pushing beyond them.
      12. Enhancing Hart and Sayegh to return top n solutions.
      13. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
      14. Adding support to abcDE for annotating phrase segmentation.
      15. Debugging Dactylize 88-key circuit.
      16. Collecting fingering data from JB performances in Elizabethtown.
      17. Completing Dactylize II circuit.
      18. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
      19. Defining procedure for sanity test of production automatic data collector (including Beringer data).
      20. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
      21. Implementing end-to-end machine learning experiment, using Beringer abcD data.
      22. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

      Wednesday, June 13, 2018

      2018-06-13 status

      Done

      Methodology

      • Contemplated a survey to find ground truth for interchangeable digits. (Discussed briefly with AFL and JB over email.)

      Model Building

      • Implemented trigram nodes in networkx for revamped Parncutt.

        Doing

        1. Learning Cytoscape for graph visualization to help debug graph code.
        2. Fixing cost functions in Parncutt.
        3. Validating Parncutt cost functions for left hand.
        4. Adding mechanism to learn weights for Parncutt rules from training data.
        5. Reading some fingering pedagogy (C. P. E. Bach, Couperin, Rami Bar-Niv) to develop a vocabulary for talking to pianists.

        Struggling


        1. A body at rest tends to stay at rest.
        2. Does this topic make sense for my new career situation?


        In Scope

        1. Implementing crude automatic segmenter.
        2. Developing staccato/legato classifier.
        3. Demonstrating improved Parncutt via #1 and #2.
        4. Debugging Sayegh model, which produces results inconsistent with training data.
        5. Developing better test cases for Sayegh.
        6. Updating abcDE to support manual segmentation.
        7. Completing and polishing abcD for entire Beringer corpus.
        8. Defining initial benchmark corpora and evaluation methodology.
        9. Implementing convenience methods for reporting benchmark results.
        10. Moving Beringer corpus to MySQL database.
        11. Enhancing Parncutt, following published techniques and pushing beyond them.
        12. Enhancing Hart and Sayegh to return top n solutions.
        13. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
        14. Adding support to abcDE for annotating phrase segmentation.
        15. Debugging Dactylize 88-key circuit.
        16. Collecting fingering data from JB performances in Elizabethtown.
        17. Completing Dactylize II circuit.
        18. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
        19. Defining procedure for sanity test of production automatic data collector (including Beringer data).
        20. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
        21. Implementing end-to-end machine learning experiment, using Beringer abcD data.
        22. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

        Friday, March 23, 2018

        2018-03-22 status

        Done

        Administrivia

        • Requested suspension of TAA support until next January (for tax purposes).
        • Answered interview questions for College of Engineering story on Fifty for the Future award.

        Model Building

        • Completed Sayegh implementation.
        • Learned how to modify TatSu AST. 
        • Implemented "pivot alignment" evaluation method.
        • Refactored code for better reuse in new models, especially for segmenting input.
        • Fixed how first and last fingering were being constrained to ensure model preferences for second and penultimate fingerings were not ignored.
        • Drafted ISMIR abstract.
        • Drew a diagram of the subproblems in the domain.

          Doing

          1. Writing up what we have done so far.
          2. Reimplementing Parncutt model in framework using networkx.
          3. Developing better test cases for Sayegh.
          4. Implementing crude automatic segmenter.
          5. Updating abcDE to support manual segmentation.
          6. Completing and polishing abcD for entire Beringer corpus.
          7. Defining initial benchmark corpora and evaluation methodology.
          8. Implementing convenience methods for reporting benchmark results.

          Struggling

          1. Sayegh model produces results that do not seem consistent with training data provided.

          In Scope

          1. Moving Beringer corpus to MySQL database.
          2. Enhancing Parncutt, following published techniques and pushing beyond them.
          3. Enhancing Hart and Sayegh to return top n solutions.
          4. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
          5. Adding support to abcDE for annotating phrase segmentation.
          6. Debugging Dactylize 88-key circuit.
          7. Collecting fingering data from JB performances in Elizabethtown.
          8. Completing Dactylize II circuit.
          9. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
          10. Defining procedure for sanity test of production automatic data collector (including Beringer data).
          11. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
          12. Implementing end-to-end machine learning experiment, using Beringer abcD data.
          13. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

          Thursday, March 22, 2018

          2018-03-01 status

          Done

          Administrivia

          • Talked through a few challenges with the NLP Lab.
          • Did a little forum shopping. As a backup plan for ISMIR conference, the TISMIR journal is accepting submissions. 

          Model Building

          • Completed "reentry" evaluation method for strike fingers.
          • Tested edge cases for evaluation and advising methods.
          • Rejected TensorFlow for Sayegh implementation. We just need a trellis graph, a nasty for loop for training, and Viterbi.
          • Stubbed in support for phrase segmentation in modeling framework.
          • Implemented Sayegh training algorithm.
          • Implemented methods to store and recall trained models for reuse.

            Doing

            1. Implementing Sayegh trellis-graph model from scratch, using Python's networkx.
            2. Defining initial benchmark corpora and evaluation methodology.
            3. Implementing convenience methods for reporting benchmark results.
            4. Completing and polishing abcD for entire Beringer corpus.

            Struggling

            1. The otherwise slick parser module I am using (TatSu) produces an immutable AST. This is cramping my style and promises to get worse as we move along.
            2. The Parncutt code is a disaster under Python 3. Lot of rework needed here.

            In Scope

            1. Reimplementing Parncutt model in framework using networkx.
            2. Moving Beringer corpus to MySQL database.
            3. Enhancing Parncutt, following published techniques and pushing beyond them.
            4. Enhancing Hart and Sayegh to return top n solutions.
            5. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
            6. Adding support to abcDE for annotating phrase segmentation.
            7. Debugging Dactylize 88-key circuit.
            8. Collecting fingering data from JB performances in Elizabethtown.
            9. Completing Dactylize II circuit.
            10. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
            11. Defining procedure for sanity test of production automatic data collector (including Beringer data).
            12. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
            13. Implementing end-to-end machine learning experiment, using Beringer abcD data.
            14. Submitting papers to ISMIR 2018. Abstracts due March 23. Papers due March 30. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.